NLP for multilingual and low resource languages
Build and evaluate language technology for languages that most datasets and models leave behind
A taste of a lesson
I want a sentiment classifier for Amharic. Can I translate an English review dataset and train on that?
You can use it as a starting point, but be careful about relying on it. Translated data carries the topics, products and ways of complaining from the English reviews, plus translation errors, so the model learns translated Amharic rather than how people actually write. A better plan: use the translated set or a multilingual model for a first version, then collect and label a few hundred genuine Amharic posts with native speakers, and keep a separate test set of original Amharic text. Quick exercise: list two ways Amharic social posts might differ from translated English reviews.
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain what makes a language low resource in NLP terms and why models struggle
- Use multilingual models and cross lingual transfer appropriately
- Handle script, spelling, dialect and code switching issues in data
- Build small, high quality datasets and native evaluation sets with speakers
- Plan a project that respects consent, data ownership and community benefit
Lesson plan
- 1 What low resource really means Understand the data, tooling and tokenisation gaps behind weak performance. Start
- 2 Multilingual models and transfer Use shared models and cross lingual transfer, and know their limits. Start
- 3 Scripts, spelling, dialects and code switching Prepare real world text that does not follow one standard form. Start
- 4 Building data with speakers Create a small, high quality labelled dataset with native speakers. Start
- 5 Evaluation that reflects the language Build test sets from original text instead of translated English. Start
- 6 Ethics and community benefit Plan a project that the language community supports and benefits from. Start
Try asking
About this tutor
For researchers, developers, linguists and community members who want NLP to work for languages with little digital data. You will learn why low resource languages fall behind, from scarce text to tokenisers that split their words into many pieces, and how multilingual models and cross lingual transfer can help. The lessons cover script, dialect and code switching challenges, building small high quality datasets with native speakers, and creating evaluation sets that come from the language itself rather than translated English. Ethics runs through the course: consent, data ownership and making sure the community benefits. Suitable for beginners with interest and for practitioners planning a project.
Reviews
Students can review a tutor after a paid lesson. Nobody has yet.
About the teacher
Practical NLP: from tokens and embeddings to classification, translation and speech
9 tutors 470 lessons taught Sample
I teach natural language processing as a craft: turning messy text in many languages into something a model can use, and checking honestly whether the result works. I grew up switching between Arabic, French and English, and my work has been on text and speech systems that had to serve speakers of more than one language, so I notice quickly...
See Nadia's profile and tutorsMore like this
Other tutors on the same or nearby topics.