Sentiment analysis: interpreting opinions in text
Learn how classifiers estimate sentiment and how to evaluate their errors.
Read moreLearn how text is divided into tokens and why different tokenization methods suit different tasks.
Content and editorial review team
Prepared by the Kinetis editorial team to explain NLP clearly and practically.
Tokenization splits text into units a model or analysis method can process. Depending on the method, those units may be words, characters, subwords, punctuation or other symbols.
Boundaries are not always obvious. The tokenizer needs rules or a learned vocabulary, and those choices affect how the text is represented.
A simple method splits on spaces. For example, The cat eats fish becomes four word-like units, before any separate handling of punctuation.
Punctuation and contractions need explicit treatment. A method suitable for one language or task may not fit another.
Contractions, compound words and technical terms create decisions about boundaries. Consider how those decisions affect the model that receives the tokens.
Canadian applications may include English, French and other languages. Apostrophes, accents and mixed-language text need representative tests rather than assumptions based on one language.
URLs, mentions and hashtags also need appropriate handling. Decide whether they carry useful information for the task before removing or splitting them.
There is no universal tokenizer for every task. When using a pretrained model, use its matching tokenizer and expected preprocessing.
Rules can handle punctuation, numbers or specific patterns. They offer control but need careful testing against edge cases.
Subword methods represent many words using smaller vocabulary units. The actual split depends on the learned vocabulary and algorithm, not simply on linguistic syllables or prefixes.
Character-based methods use small units and can handle varied words, but often create longer sequences and different modelling trade-offs.
Established NLP libraries offer tokenizers, so you can study the behaviour of an existing method before implementing your own.
NLTK provides educational text-processing tools. Check the language resources and tokenizer behaviour for your use case.
spaCy provides language-specific tokenization and processing components. Check its documentation and test representative input.
Hugging Face Transformers includes interfaces to tokenizers associated with pretrained models. Keep tokenizer and model versions compatible.
Good tokenization deserves attention before a model is trained.
A practical NLP reminder
Tokenization affects every later stage that consumes the tokens. Inspect its output early to catch unexpected boundaries or missing information.
Choose a method that fits your model and language coverage, then test it against realistic examples. A familiar library does not remove the need to inspect the result.
Start with a short sample containing punctuation, accents, contractions and domain terms. Compare the output with the representation your application needs.
The appropriate tokenizer depends on the model, language and task. This guide explains general concepts; validate implementation choices before production use.
Learn how classifiers estimate sentiment and how to evaluate their errors.
Read more
Explore intent handling, response generation and dialogue flow.
Read more
Understand the foundations of language models and Transformers in clear language.
Read more