SophiArch
PractitionerML202

Natural Language Processing with Python

Understand how text becomes data: tokenisation, representations, embeddings, and evaluation. Built around the judgment skills practitioners need to direct AI tools confidently and catch silent NLP failures.

Lessons are AI-assisted and human-reviewed. Learn more.

Syllabus

Text as Data

01
The NLP Workflow: What "Correct" Looks LikeFree preview
30 min
02
Tokenisation and Normalisation: What AI Tools Change Without Telling You
35 min
03
Text Representations: From Counts to Embeddings
40 min

Classical NLP as a Baseline

04
TF-IDF and Document Similarity: The Baseline You Must Know
35 min
05
Text Classification Pipelines: Building Them Right (and Catching What AI Gets Wrong)
42 min

LLM-Native NLP

06
LLM-Assisted NLP: When to Prompt, When to Train
35 min

Embeddings and Semantic Understanding

07
Word and Sentence Embeddings: What They Encode, What They Miss
40 min
08
Semantic Similarity and Vector Search: The Foundation of RAG
38 min
09
Embedding Failure Modes: What AI Gets Wrong and How to Catch It
35 min

Evaluation and Production

10
Evaluating NLP Models: Metrics Beyond Accuracy
40 min
11
End-to-End NLP Pipelines: From Raw Text to Production
42 min