CTU-AI323.AJ1
Natural Language Processing
- Practice in 32 Hands-On Labs — nothing to install
- 11 Interactive Lessons and 283 topics mapped to the official exam objectives
Expert Self-paced · 1 year access
32 Hands-On LiveLabs
Practice real IT tasks in guided environments.
- Real environments
- Auto-graded
- No installation
11Interactive Lessons
283Topics
32LiveLab
4Videos
73Flashcards
73Glossary of terms
01 / Lessons & labs
See exactly what you will learn and practice
Lessons
11 Interactive Lessons · 283 topics01 Statistical Foundations of NLP 38 topics · 8 LiveLab +
- What is NumPy?
- What are NumPy Arrays?
- Lists and Exponents
- Arrays and Exponents
- Calculating the Mean and Standard Deviation
- Plotting a Line with NumPy and Matplotlib
- What is Linear Regression?
- The MSE Formula
- What is Pandas?
- A Pandas Data Frame with NumPy Example
- Describing a Pandas Data Frame
- Reading CSV Files in Pandas
- The loc() and iloc() Methods in Pandas
- Converting Categorical Data to Numeric Data
- Combining Pandas Data frames
- Data Manipulation with Pandas Data Frames (1)
- Data Manipulation with Pandas Data Frames (2)
- Handling Missing Data in Pandas
- Sorting Data Frames in Pandas
- Working with groupby() in Pandas
- Pandas Data Frames and Simple Statistics
- Aggregate Operations in Pandas Data Frames
- Working with JSON-based Data
- What is Text Encoding?
- Text Encoding Techniques
- The BoW Algorithm
- What are n-grams?
- Calculating tf, idf, and tf-idf
- The Context of Words in a Document
- What is Cosine Similarity?
- Text Vectorization (aka Word Embeddings)
- Overview of Word Embeddings and Algorithms
- What is Word2vec?
- The CBoW Architecture
- What are Skip-grams?
- What is GloVe?
- Comparison of Word Embeddings
- Language Models and NLP
8 LiveLab in this lesson — see the labs panel →
02 Machine Learning for NLP Tasks 37 topics · 6 LiveLab +
- Cleaning Data with Regular Expressions
- Handling Contracted Words
- Python Code Samples of BoW
- One-Hot Encoding Examples
- Sklearn and Word Embedding Examples
- Web Scraping with Pure Regular Expressions
- What is SpaCy?
- What is NLTK?
- NLTK and BoW
- NLTK and Stemmers
- NLTK and Lemmatization
- NLTK and Stop Words
- What is Wordnet?
- NLTK and n-grams
- NLTK and POS (1)
- NLTK and POS (2)
- NLTK and Tokenizers
- What is Gensim?
- An Example of Topic Modeling
- A Brief Comparison of Popular Python-Based NLP Libraries
- What is Classification?
- What are Linear Classifiers?
- What is kNN?
- What are Decision Trees?
- Decision Trees, Gini Impurity, and Entropy
- What are Random Forests?
- What are Support Vector Machines?
- What is a Bayesian Classifier?
- Training Classifiers
- Evaluating Classifiers
- Trade-offs for ML Algorithms
- What are Activation Functions?
- Common Activation Functions
- The ReLU and ELU Activation Functions
- Sigmoid, Softmax, and Hardmax Similarities
- Hyperparameters for Neural Networks
- What is Logistic Regression?
6 LiveLab in this lesson — see the labs panel →
03 NLP Applications Across Domains 30 topics · 6 LiveLab +
- What is Machine Learning?
- Types of Machine Learning Algorithms
- Preparing a Dataset and Training a Model
- Feature Engineering, Selection, and Extraction
- Working with Datasets
- Overfitting versus Underfitting
- Data Normalization Techniques
- Metrics in Machine Learning
- What is Linear Regression?
- Other Types of Regression
- The Mean Squared Error (MSE) Formula
- Calculating the MSE Manually
- What are Ensemble Methods?
- Four Types of Ensemble Methods
- Common Boosting Algorithms
- Hyperparameter Optimization
- AutoML, AutoML-Zero, and AutoNLP
- What is Text Summarization?
- Text Summarization with gensim and SpaCy
- What are Recommender Systems?
- Content-Based Recommendation Systems
- Collaborative Filtering Algorithm
- What is Sentiment Analysis?
- Sentiment Analysis with Naïve Bayes
- Sentiment Analysis in NLTK and VADER
- Sentiment Analysis with Textblob
- Sentiment Analysis with Flair
- Detecting Spam
- Logistic Regression and Sentiment Analysis
- What are Chatbots?
6 LiveLab in this lesson — see the labs panel →
04 Performance Optimization in NLP Systems 17 topics · 6 LiveLab +
- Term-Document Matrix
- Text Classification Algorithms in Machine Learning
- A Keras-Based Tokenizer
- TF2 and Tokenization
- TF2 and Encoding
- A Keras-Based Word Embedding
- An Example of BoW with TF2
- The 20newsgroup Dataset
- Text Classification with the kNN Algorithm
- Text Classification with a Decision Tree Algorithm
- Text Classification with a Random Forest Algorithm
- Text Classification with the SVC Algorithm
- Text Classification with the Naïve Bayes Algorithm
- Text Classification with the kMeans Algorithm
- TF2/Keras and Word Tokenization
- TF2/Keras and Word Encodings
- Text Summarization with TF2/Keras and Reuters Dataset
6 LiveLab in this lesson — see the labs panel →
05 Research and Emerging Trends in NLP 15 topics · 6 LiveLab +
- What is Attention?
- An Overview of the Transformer Architecture
- What is T5?
- What is BERT?
- The Inner Workings of BERT
- Subword Tokenization
- Sentence Similarity in BERT
- Generating BERT Tokens (1)
- Generating BERT Tokens (2)
- The BERT Family
- Introduction to GPT
- Working with GPT-2
- What is GPT-3?
- The Switch Transformer: One Trillion Parameters
- Looking Ahead
6 LiveLab in this lesson — see the labs panel →
06 Appendix A: Data and Statistics 24 topics +
- What are Datasets?
- Preparing Datasets
- Missing Data, Anomalies, and Outliers
- What is Imbalanced Classification?
- What is SMOTE?
- Analyzing Classifiers
- What is a Probability?
- Random Variables
- Fundamental Concepts in Statistics
- The Moments of a Function (Optional)
- Data and Statistics
- The Bias-Variance Trade-off
- Gini Impurity, Entropy, and Perplexity
- Cross-Entropy and KL Divergence
- Covariance and Correlation Matrices
- Principal Component Analysis (PCA)
- Dimensionality Reduction
- Dimensionality Reduction Techniques
- Linear Versus Nonlinear Reduction Techniques
- Types of Distance Metrics
- Other Well-Known Distance Metrics
- What is Sklearn?
- What is Bayesian Inference?
- What are Vector Spaces?
07 Appendix B: Introduction to Python 28 topics +
- Tools for Python
- Python Installation
- Setting the PATH Environment Variable (Windows Only)
- Launching Python on Your Machine
- Python Identifiers
- Lines, Indentation, and Multilines
- Quotation and Comments in Python
- Saving Your Code in a Module
- Some Standard Modules in Python
- The help() and dir() Functions
- Compile Time and Runtime Code Checking
- Simple Data Types in Python
- Working with Numbers
- Working with Fractions
- Unicode and UTF-8
- Working with Unicode
- Working with Strings
- Uninitialized Variables and the Value None in Python
- Slicing and Splicing Strings
- Search and Replace a String in Other Strings
- Remove Leading and Trailing Characters
- Printing Text without NewLine Characters
- Text Alignment
- Working with Dates
- Exception Handling in Python
- Handling User Input
- Python and Emojis (Optional)
- Command-Line Arguments
08 Appendix C: Introduction to Regular Expressions 23 topics +
- What are Regular Expressions?
- Metacharacters in Python
- Character Sets in Python
- Character Classes in Python
- Matching Character Classes with the re Module
- Using the re.match() Method
- Options for the re.match() Method
- Matching Character Classes with the re.search() Method
- Matching Character Classes with the findAll() Method
- Additional Matching Function for Regular Expressions
- Grouping with Character Classes in Regular Expressions
- Using Character Classes in Regular Expressions
- Modifying Text Strings with the re Module
- Splitting Text Strings with the re.split() Method
- Splitting Text Strings Using Digits and Delimiters
- Substituting Text Strings with the re.sub() Method
- Matching the Beginning and the End of Text Strings
- Compilation Flags
- Compound Regular Expressions
- Counting Character Types in a String
- Regular Expressions and Grouping
- Simple String Matches
- Additional Topics for Regular Expressions
09 Appendix D: Introduction to Keras 10 topics +
- What is Keras?
- Creating a Keras-Based Model
- Keras and Linear Regression
- Keras, MLPs, and MNIST
- Keras, CNNs, and cifar10
- Resizing Images in Keras
- Keras and Early Stopping (1)
- Keras and Early Stopping (2)
- Keras and Metrics
- Saving and Restoring Keras Models
10 Appendix E: Introduction to TensorFlow 2 31 topics +
- What is TF 2?
- Other TF 2-Based Toolkits
- TF 2 Eager Execution
- TF 2 Tensors, Data Types, and Primitive Types
- Constants in TF 2
- Variables in TF 2
- The tf.rank() API
- The tf.shape() API
- Variables in TF 2 (Revisited)
- What is @tf.function in TF 2?
- Working with @tf.function in TF 2
- Arithmetic Operations in TF 2
- Caveats for Arithmetic Operations in TF 2
- TF 2 and Built-In Functions
- Calculating Trigonometric Values in TF 2
- Calculating Exponential Values in TF 2
- Working with Strings in TF 2
- Working with Tensors and Operations in TF 2
- Second-Order Tensors in TF 2 (1)
- Second-Order Tensors in TF 2 (2)
- Multiplying Two Second-Order Tensors in TF
- Convert Python Arrays to TF Tensors
- Differentiation and tf.GradientTape in TF 2
- Examples of tf.GradientTape
- What is Trax?
- Google Colaboratory
- Other Cloud Platforms
- TF2 and tf.data.Dataset
- The TF 2 tf.data.Dataset
- What are Lambda Expressions?
- Working with Generators in TF 2
11 Appendix F: Data Visualization 30 topics +
- What is Data Visualization?
- What is Matplotlib?
- Horizontal Lines in Matplotlib
- Slanted Lines in Matplotlib
- Parallel Slanted Lines in Matplotlib
- A Grid of Points in Matplotlib
- A Dotted Grid in Matplotlib
- Lines in a Grid in Matplotlib
- A Colored Grid in Matplotlib
- A Colored Square in an Unlabeled Grid in Matplotlib
- Randomized Data Points in Matplotlib
- A Histogram in Matplotlib
- A Set of Line Segments in Matplotlib
- Plotting Multiple Lines in Matplotlib
- Trigonometric Functions in Matplotlib
- Display IQ Scores in Matplotlib
- Plot a Best-Fitting Line in Matplotlib
- Introduction to Sklearn (scikit-learn)
- The Digits Dataset in Sklearn
- The Iris Dataset in Sklearn
- The Iris Dataset in Sklearn (Optional)
- The faces Dataset in Sklearn (Optional)
- Working with Seaborn
- Seaborn Built-in Datasets
- The Iris Dataset in Seaborn
- The Titanic Dataset in Seaborn
- Extracting Data from the Titanic Dataset in Seaborn (1)
- Extracting Data from the Titanic Dataset in Seaborn (2)
- Visualizing a Pandas Dataset in Seaborn
- Data Visualization in Pandas
Hands-On Labs Our edge
32 LiveLabs- Using Lists and Arrays
- Performing Statistical Operations
- Creating Line Charts
- Performing Linear Regression
- Creating and Accessing DataFrames
- Working with Data Frames - I
- Working with Data Frames - II
- Analyzing JSON Data Structures
- Cleaning Data and Handling Contracted Words
- Analyzing Encoding and Embedding Techniques
- Performing Web Scraping
- Using Gensim for Text Processing
- Using NLTK for Text Processing
- Selecting the Best Classification Model
- Performing Data Normalization Using scikit-learn
- Plotting and Interpreting ROC Curves and AUC Scores
- Performing Sentiment Analysis Using NLTK and the Multinomial Naïve Bayes Classifier
- Evaluating Regression Models Using Error Metrics
- Summarizing Text with Gensim and spaCy
- Creating a Spam Classifier
- Preparing Text Data for Deep Learning-Based NLP Models
- Creating Word Embeddings Using Keras Neural Networks for Text Classification
- Implementing BoW Text Representation Using Tensorflow/Keras
- Performing Text Processing Using TF2 and Keras
- Comparing Text Classification Models
- Performing Text Tokenization and N-gram Extraction Using TensorFlow Text
- Using the transformers Library
- Performing NER Using Hugging Face Transformers
- Performing Subword Tokenization Using WordPiece And BPE
- Generating BERT Tokens
- Performing Text Generation Using GPT-2
- Analyzing GPT Models
Labs run in your browser — nothing to install.
02 / FAQs
Questions before you start
Prepare for Natural Language Processing
One-time payment. Full access for 1 year. Start with a free trial if you want to look around first.
- 1 year of full access
- 32 LiveLab included
- Certificate of completion
No credit card required