From the archive

SEA-LION v3 builds regional language support on Gemma 2

AI Singapore adapts a nine-billion-parameter model with Southeast Asian training data and regional instruction tuning.

By Chat Overview Published Updated

AI Singapore introduced SEA-LION v3, a language model built by continuing the training of Google's Gemma 2 9B. The work uses 200 billion tokens covering eleven official Southeast Asian languages.

Regional language and context

The release combines continued training with instruction tuning, model merging, and alignment. Project SEALD, a partnership involving AI Singapore and Google, contributed regional data and technical support.

The aim is to improve how a model handles the region's languages and cultural context, rather than relying only on a broadly trained model's English-language strengths.

A model to adapt and deploy

The nine-billion-parameter size gives the project a relatively compact base for further development. It follows the earlier Llama-based SEA-LION v2 and extends the project's work across model families. The SEA-LION profile covers available models, access routes, and the terms associated with each release.

Source