Open research infrastructure for Afaan Oromoo AI
Building OromoCorpus, OromoTokenizer, OromoLM, and OromoBench.
π Explore the Oromo AI Research Hub
Browse our research, documentation, corpus work, tokenizer experiments, model roadmap, and project progress as an interactive website.
OromoCorpus β OromoTokenizer β OromoLM β OromoBench β Applications
Oromo AI is an independent open-source research and engineering project focused on building high-quality AI infrastructure for Afaan Oromoo.
The project follows a data-first path:
licensed sources
β
provenance + audit
β
cleaning + deduplication
β
OromoCorpus
β
tokenizer research
β
OromoLM continued pretraining
β
OromoBench
β
instruction tuning + applications
The goal is not to build a quick chatbot. The goal is to create reusable, documented, and reproducible infrastructure for Afaan Oromoo language modeling, translation, retrieval, evaluation, and future speech systems.
| Component | Purpose | Status |
|---|---|---|
| OromoCorpus | Versioned, provenance-tracked Afaan Oromoo training corpus | β 50M minimum achieved; expanding toward 100M |
| OromoTokenizer | Tokenization research and Oromo-aware tokenizer candidates | β Gemma +8K training candidate validated |
| OromoLM | Continued-pretrained Afaan Oromoo causal language models | π§ͺ Exact three-way Oromo BPB evaluation complete |
| OromoBench | Afaan Oromoo evaluation framework | π In development |
Canonical naming is documented in docs/NAMING.md.
The current accepted corpus planning total is measured with the Oromo Unigram 48K + byte-fallback tokenizer as a research reference. It is not yet the final OromoLM tokenizer.
| Accepted source | Net-new records | 48K reference tokens |
|---|---|---|
| AfriBERTa Afaan Oromoo v0.1.2 | 410,193 | 9,587,934 |
| Wikimedia omwiki v0.1 | 2,254 | 1,070,896 |
| VOA Afaan Oromoo via WURA v0.1 | 9,510 | 1,899,811 |
| WaxalNLP Oromo ASR v0.1 | 44,194 | 1,934,045 |
| MADLAD-400 Oromo v0.2 | 18,704 | 17,873,390 |
| HPLT 3.0 gaz_Latn v0.1 | 26,655 | 19,771,727 |
| Accepted total | 511,510 | 52,137,803 |
MADLAD-400 Oromo v0.2 is now accepted into OromoCorpus under the upstream AllenAI MADLAD-400 dataset's published ODC-BY license. The approved frozen research-clean artifact contains 18,704 records / 17,873,390 48K-reference tokens.
This approval relies on the upstream dataset-level license representation and
preserves attribution to AllenAI / MADLAD-400. It does not claim that Oromo
AI independently cleared copyright for every underlying Common Crawl page.
That scope limitation, along with the completed provenance investigation, is
permanently documented in
MADLAD_400_LICENSE_DECISION.md.
Official v1.5 provenance was independently recovered for 1,917 final records / 1,874,398 reference tokens (10.49%) across 182 domains. Those findings remain part of the audit trail even though the full frozen subset is accepted under the upstream dataset license.
HPLT 3.0 gaz_Latn v0.1 is now accepted/frozen as the sixth OromoCorpus
source. The approved subset is restricted to WDS bins 8β10, passed exact
and canonical near-deduplication against all previously accepted sources,
passed structural review, and was fully checked with GlotLID v3. The final
artifact contains 26,655 records / 19,771,727 48K-reference tokens across
1,091 unique source domains.
HPLT publishes the dataset packaging under CC0, while explicitly stating
that it does not own the underlying extracted text. Oromo AI therefore records
the underlying individual-content rights as not independently verified and
does not represent every source webpage as CC0. See
HPLT3_LICENSE_DECISION.md.
OromoCorpus v0.2 minimum: 50,000,000 tokens
Current planning total: 52,137,803
Progress: 104.28%
Margin above minimum: 2,137,803
OromoCorpus v0.3 target: 100,000,000 tokens
Current progress: 52.14%
The full WURA Oromo package remains under source-level review and does not count as an accepted source by itself. Only independently audited and rights-cleared subsets are admitted.
With the 50M minimum now achieved, corpus acquisition continues toward the preferred 100M target with greater emphasis on domain, dialect, literary, educational, technical, and conversational diversity rather than raw volume alone.
Tokenizer research has established:
| Tokenizer | Vocabulary | Tokens/word | Fragmentation | UNK |
|---|---|---|---|---|
| Oromo Unigram 48K + byte fallback | 48,000 | 1.4000 | 38.75% | 0 |
| Oromo Unigram 32K + byte fallback | 32,000 | 1.4495 | 40.83% | 0 |
| AfriBERTa | 70,006 | 1.6765 | 39.83% | 0 |
Native causal-model tokenizers were also benchmarked on the frozen Oromo evaluation set. Whole-word vocabulary augmentation improved fragmentation substantially, but it did not match the custom Oromo tokenizer's sequence efficiency.
Full results: docs/TOKENIZER_RESEARCH_REPORT.md.
The first controlled model-level CPT comparison is complete on google/gemma-3-1b-pt, comparing the native tokenizer against the frozen Gemma +8K Afaan Oromoo whole-word augmentation on the same six-source pilot text.
| Metric | Native Gemma | Oromo +8K | Change |
|---|---|---|---|
| Training tokens | 5,708,421 | 4,811,939 | -15.70% |
| Optimizer steps | 349 | 294 | -15.76% |
| Train wall time | 1,437.05 s | 1,176.71 s | -18.12% |
| Training throughput | 3,972.59 tok/s | 4,090.03 tok/s | +2.96% |
| Peak reserved VRAM | 12.21 GiB | 12.55 GiB | +2.78% |
The comparison ran on the same NVIDIA A100-SXM4-80GB environment with sequence length 1024, BF16, gradient accumulation 16, and one epoch of identical underlying text exposure.
The efficiency result remains positive, but final tokenizer selection remains open. Exact document-reset, byte-normalized evaluation is now complete on the same frozen 1,503-record Oromo validation set:
| Model state | Exact BPB | Relative to base |
|---|---|---|
Original google/gemma-3-1b-pt |
2.670046 | baseline |
| Native tokenizer + Oromo CPT | 1.733832 | -35.06% |
| OromoLM +8K + Oromo CPT | 1.863339 | -30.21% |
Lower BPB is better. This establishes that OromoCorpus continued pretraining substantially improves held-out Afaan Oromoo modeling. The +8K candidate also improves strongly over base Gemma, but after one epoch it remains about 7.47% higher/worse in BPB than native-tokenizer CPT.
The next gate is controlled longer +8K CPT with exact BPB checkpoints and a frozen general-language forgetting control. The current native-CPT quality benchmark to beat is 1.733832 BPB.
Full experiment records: docs/CPT_PILOT_REPORT.md and docs/CPT_RERUN_V0_2_RESULTS.md.
At the latest verified checkpoint:
104 tests passed
The test suite covers corpus schema, ingestion, cleaning, quality decisions, deduplication, validation, and source-registry behavior.
Oromo AI uses a provenance-first corpus policy.
Every accepted source should have:
- identifiable origin and ownership;
- an explicit license or training-use decision;
- immutable raw-source handling;
- conservative cleaning;
- exact and near-duplicate removal;
- reproducible hashes and manifests;
- source-level acceptance/rejection records;
- evaluation-data separation.
The pipeline intentionally avoids destructive normalization that could erase Qubee orthography, dialectal variation, morphology, punctuation, or legitimate multilingual context.
Detailed policy: docs/CLEANING_POLICY.md.
oromo-ai/
βββ data/ # source registry, manifests, processed corpus metadata
βββ docs/ # research reports, policies, roadmap, source audits
βββ reports/ # generated research/evaluation outputs
βββ scripts/ # reproducible corpus and research utilities
βββ src/ # Python package and data pipeline
βββ tests/ # automated tests
βββ tokenizer/ # tokenizer training and evaluation artifacts
βββ pyproject.toml
βββ uv.lock
βββ README.md
Generated raw/interim corpus artifacts are intentionally kept out of Git where appropriate.
Requirements:
- Python 3.12
- uv
Install dependencies:
uv syncRun the test suite:
uv run pytest -qInspect the source registry:
uv run python scripts/source_registry.py list
uv run python scripts/source_registry.py validateThe repository is research infrastructure under active development; commands and artifact formats may evolve as corpus and model work progresses.
β
Corpus pipeline and provenance framework
β
AfriBERTa Oromo v0.1.2 validated
β
Wikimedia omwiki source qualified
β
VOA Afaan Oromoo subset qualified
β
WaxalNLP Oromo ASR source qualified
β
MADLAD-400 Oromo technical and quality qualification complete
β
MADLAD-400 Oromo provenance recovery: 1,917 records / 1,874,398 tokens mapped
β
MADLAD-400 Oromo v0.2 approved under upstream ODC-BY dataset license
β
Frozen tokenizer evaluation set
β
Custom tokenizer benchmark
β
Native causal-tokenizer benchmark
β
Whole-word augmentation study
β
Cleaned Gemma +8K tokenizer frozen as a training candidate
β
Dual-tokenizer workload audit and 1024-token packing
β
Native-vs-+8K 10-step GPU smoke tests
β
First controlled Gemma native-vs-+8K CPT pilot
β
OromoCorpus 50M minimum achieved β 52,137,803 reference tokens
π Expand OromoCorpus toward the preferred 100M target
β
Exact byte-normalized three-way LM evaluation
π Longer +8K CPT learning curve + forgetting controls
π Develop OromoBench
β
Tiny CPT proof
β³ Scaled continued pretraining
β³ Supervised instruction tuning
β³ Model release engineering
β³ Translation / retrieval / speech applications
The complete roadmap is maintained in docs/ROADMAP.md.
Detailed technical material lives in docs/ rather than being duplicated in this README.
| Document | Purpose |
|---|---|
| Corpus Expansion Plan | 50M/100M acquisition and release strategy |
| Corpus Statistics Report | Corpus measurements and validation |
| Corpus Quality Report | Quality analysis |
| Cleaning Policy | Conservative preprocessing rules |
| Tokenizer Research Report | Tokenizer benchmarks and experiments |
| CPT Pilot Report | Gemma native-vs-+8K workload, GPU, quality, and reproducibility evidence |
| CPT Reproduction v0.2 Results | Exact three-way Oromo BPB results, rerun hashes, and next quality gate |
| Evaluation | OromoBench and normalized model-evaluation direction |
| Roadmap | Project phases and current milestone |
| MADLAD License Decision | ODC-BY approval basis, attribution obligations, and scope limitations |
| MADLAD Provenance Review | Partial source-level provenance recovery, URL/domain audit, and VOA review |
| HPLT3 Oromo Report | HPLT3 filtering, deduplication, quality, language verification, and frozen metrics |
| HPLT3 License Decision | CC0 packaging scope, underlying-text caveat, and project acceptance basis |
| Naming | Canonical project naming |
Source-specific qualification reports are maintained under docs/sources/.
Useful contributions include:
- rights-clear Afaan Oromoo corpora;
- linguistic review and annotations;
- dialect and domain metadata;
- parallel corpora;
- tokenizer and evaluation research;
- benchmark construction;
- reproducible data-engineering improvements.
All contributed data must preserve clear provenance and licensing information.
Data before model. Evidence before scale.
Oromo AI is being built one validated layer at a time.