From the archive
SEA-LION v3 builds regional language support on Gemma 2
AI Singapore adapts a nine-billion-parameter model with Southeast Asian training data and regional instruction tuning.
AI Singapore introduced SEA-LION v3, a language model built by continuing the training of Google's Gemma 2 9B. The work uses 200 billion tokens covering eleven official Southeast Asian languages.
Regional language and context
The release combines continued training with instruction tuning, model merging, and alignment. Project SEALD, a partnership involving AI Singapore and Google, contributed regional data and technical support.
The aim is to improve how a model handles the region's languages and cultural context, rather than relying only on a broadly trained model's English-language strengths.
A model to adapt and deploy
The nine-billion-parameter size gives the project a relatively compact base for further development. It follows the earlier Llama-based SEA-LION v2 and extends the project's work across model families. The SEA-LION profile covers available models, access routes, and the terms associated with each release.