AI Research Lab · Project

Native Indian LLM

A sovereign foundation large language model, designed and trained in India.

← Back to AI Lab

Overview

Our goal is a frontier-class model that understands Indian languages, scripts, and code-mixing out of the box — reducing dependence on foreign models and giving Indian developers and enterprises a model they fully own: the weights, the data pipeline, and the roadmap.

First model

Saffron-v1 (~100M)

The first concrete step is Saffron-v1 — a compact, efficient ~100M-parameter model with our own architecture (RoPE, RMSNorm, SwiGLU, grouped-query attention, and QK-normalization) and a custom byte-level BPE tokenizer, trained on a curriculum English corpus (TinyStories → Wikipedia → FineWeb-Edu). A first Preliminary ~100M checkpoint has been trained and published — an Experimental foundation we build on toward broader Indian-language coverage. An instruction-tuned chat variant (Experimental) is also published; being ~100M and lightly tuned, it follows the chat format but can produce inaccurate answers.

Saffron-v1 on GitHub Model on Hugging Face Chat model (Experimental)

Objectives

  • Language coverageStrong performance across major Indian languages and English.
  • SovereigntyFully owned weights, data, and training pipeline.
  • PracticalityReady for chat, RAG, coding, and agentic applications.
  • Safety & alignmentCulturally aware, safe, and reliable outputs.

Our Approach

  • ArchitectureOur own ~100M decoder: RoPE, RMSNorm, SwiGLU, GQA, and QK-norm.
  • TokenizerCustom byte-level BPE — frequent words become single tokens, zero out-of-vocabulary.
  • DataCurriculum corpus — TinyStories → Wikipedia → FineWeb-Edu (~1B tokens prepared), scaling toward Indian languages.
  • TrainingPretraining first; instruction tuning and preference optimization to follow.

Focus & Tech

Own ArchitectureCustom BPECurriculum DataEfficient (~100M)Edge-friendly

Status: Experimental — first Preliminary checkpoint trained (~100M params, ~500M training tokens, bf16 on GPU).

  • Preliminary resultValidation loss 3.46 (perplexity ≈ 31.9) on a held-out split. Early proof-of-life — undertrained by design and not instruction-tuned; outputs may be inaccurate.
View Code on GitHub Model on Hugging Face

Interested in this work?

We collaborate on foundation-model research for India.

Get in Touch