AI & ML interests
None defined yet.
Recent Activity
Articles
LILT
We build the multilingual layer for English-first AI. Custom evals, benchmarks, and RL environments across 200+ languages.
Most agent and coding benchmarks ship in English. We build non-English counterparts grounded in language and culture, along with the multilingual environments models train on, so labs and enterprises can measure and improve how their models perform in the languages their users speak.
New: AURORA, the multilingual AI leaderboard
AURORA ranks frontier models on non-English agentic tasks grounded in language and culture. Every task is built or verified by native-language domain experts, so results reflect how models perform for real users, not how well they handle a translation. Benchmarks at launch, with more coming:
- Multilingual Terminal-Bench: agentic coding, 324 tasks across AR / CS / DE / ES / HI / JA / KO / SR / TR / ZH
- Multilingual τ³-bench: multi-turn customer support with tool use (airline, retail, telecom, banking), pass^1 and pass^4
- Multilingual MultiChallenge: long-context instruction following, memory and self-coherence
- GAIA-v2-LILT: agentic reasoning and tool use across AR / DE / HI / KO / PT-BR
Harness, trial counts, reasoning settings and confidence intervals are published for every run. See the AURORA collection below.
Why we publish here
Open releases make it easier for the community to stress-test our work, reproduce our scores, and extend our benchmarks to new languages. We publish each benchmark with its methodology and explicit limitations.
What you'll find here
- Benchmarks & datasets: multilingual evaluations across coding, agents, tool use, long context and instruction following
- Baselines: frontier-model scores with exact setup and dated snapshots, tracked on AURORA
- Articles & papers: methodology, audit workflow, and findings
Links
- AURORA leaderboard: https://aurora.lilt.com
- Discord: https://discord.com/invite/KQt26hGZtY
- Website: https://lilt.com
- GitHub: https://github.com/lilt
- X: https://x.com/LILTLabs
- Custom and private benchmarks: https://lilt.com/contact/ai-data-services
Citation
If you use one of our datasets or benchmarks, please cite the paper linked on its dataset card or article.