Tokenization: turning text into useful units
Learn how text is divided into tokens and why different tokenization methods suit different tasks.
Foundations, methods and architectures for developers in Canada
2026 learning resources
Explore NLP, sentiment analysis and language models
Learn how text is divided into tokens and why different tokenization methods suit different tasks.
Explore sentiment classification, labelled data and evaluation, including the challenges of ambiguous or mixed opinions.
Understand intent handling, context, storage and dialogue flow in a conversational application.
Follow the ideas behind modern language models and learn when adaptation may be useful.
Build understanding in stages, from text preparation to model-based applications
Explore how text becomes input for an analysis pipeline. Consider normalization and special characters carefully: removing words or punctuation is not appropriate for every model or task.
Use labelled examples to build a simple baseline, then evaluate its predictions on held-out data. Look at errors as well as an overall score.
Combine intent handling and conversation state in a limited use case. Plan how the system should respond when it does not understand a request.
Move on to attention, Transformers and model adaptation. Examine evaluation, resource requirements and deployment constraints alongside model capabilities.