Home›Materials›AI & Data Science›Natural Language Processing – Unit IV: Language Modeling
About this material
Complete Natural Language Processing Unit IV notes covering Language Modeling and N-Gram Models. The material explains language models, unigrams, bigrams, trigrams, probabilistic language models, the Chain Rule of Probability, N-Gram and Markov Models, and Maximum Likelihood Estimation (MLE).
Topics also include Language Model Evaluation using Perplexity, Accuracy, F1 Score, BLEU Score, ROUGE Score and Human Evaluation; Parameter Estimation and Smoothing; Bayesian Parameter Estimation; Language Model Adaptation; different types of language models including N-Gram, Neural, Transformer-based, Contextualized, Encoder-Decoder, Hybrid and Pre-trained Language Models.
The unit further covers language-specific modeling problems, multilingual and cross-lingual language modeling, multilingual embeddings, machine translation, zero-shot and few-shot learning, and cross-lingual Named Entity Recognition (NER).