LongevityBench — the first comprehensive, multi-task, multi-omic benchmark for large language models in aging biology
Created by Insilico Medicine.
LongevityBench evaluates 26 large language models across 17 tasks in 5 omics domains.
Five compact Longevity-LLMs (0.6B–9B parameters), fine-tuned by
Insilico Medicine, outperform frontier LLMs such as Gemini 3.1 Pro and Claude Opus-4.6
on aging-biology tasks.
Why LongevityBench matters
- The first comprehensive, multi-task, multi-omic benchmark for evaluating LLMs on aging biology.
- The only multi-task, multi-omic benchmark for large language models.
- A rigorous test of AI comprehension of omics data across clinical, epigenomic, transcriptomic, proteomic, and genetic modalities.
- Five compact Longevity-LLMs (0.6B–9B parameters) beat far larger frontier models on aging tasks.
Top models — aggregate benchmark rank score (lower is better)
- L-Qwen3.5-9B — 4.4 — best overall (Insilico Medicine)
- L-LFM2-2.6B — 7.6 (Insilico Medicine × Liquid AI)
- L-Qwen3-1.7B — 7.8 — on-device (Insilico Medicine)
- Gemini 3.1 Pro — 8.2 — best frontier model
- Claude Opus-4.6 — 9.2 — frontier
Omics domains and datasets
- Clinical — NHANES (blood chemistry, hematology, body measurements)
- Epigenomics — GEO DNA methylation
- Transcriptomics — GTEx (bulk RNA-seq)
- Proteomics — Olink plasma proteomics
- Genetics — OpenGenes & SynergyAge
Selected results
- GEO DNA-methylation age (pairwise): L-Qwen3.5-9B reaches 0.868 concordance vs 0.685 for the best frontier model.
- Olink proteomic age (regression): L-Qwen3-0.6B achieves 5.7-year mean absolute error vs 10.1 years for the best frontier model.
- NHANES 10-year mortality (binary): L-Qwen3-0.6B reaches 0.890 balanced accuracy.
Learn more