Nia 1.0
Frontier medical intelligence, built in the UAE.
A 28-billion-parameter multimodal healthcare language model — clinical reasoning, medical examinations, literature and imaging, from a model small enough to run on a single machine, entirely inside your own infrastructure.
- parameters
- 28Bparameters
- multimodal
- Text + imagemultimodal
- context
- 262Kcontext
Weights are available to qualified organisations under an evaluation agreement. Nia is a clinical decision support tool — not a licensed clinician, and not a substitute for professional medical judgement.
Ahead of every medical model with a published figure — including Google’s MedGemma and M42’s Med42.
- MedGemma 27BGoogle87.7
- Med42-v2 70BM42 · Abu Dhabi87.3
- Med-PaLM 2Google86.5
Ahead of the best published text results — DeepSeek-R1 37.8 and o3-mini 37.3, per the official leaderboard.
Nia measured in bfloat16, single greedy pass, temperature 0, zero-shot. Competitor figures are published by their vendors.
Benchmarked against
- MedGemma 27B
- Med42-v2 70B
- Med-PaLM 2
- MedGemma 1.5 4B
- GPT-5
- GPT-5.4
- Gemini 2.5 Pro
- Claude Opus 4.5
- DeepSeek-V3.1
- DeepSeek R2
- Qwen 3.5
- Llama 4 Maverick
- DeepSeek-R1
- o3-mini
- o1
- GPT-4o
- Gemini-2.0-Flash
Benchmarks
Measured against the models built for medicine
The comparison that matters most is against other healthcare models. Nia clears MedGemma 27B’s published figures on 4 of 5 text benchmarks — by up to +15.3 points — and is ahead of Google’s Med-PaLM 2 and M42’s Med42-v2 on MedQA. PubMedQA is the one it loses, and it is on this page for that reason.
MedQA
93.0
+5.3vs MedGemmaUSMLE-style medical licensing exam questions
MedMCQA
76.5
+2.3vs MedGemmaIndian medical entrance exam, multi-option
PubMedQA
71.5
−5.3vs MedGemmaYes/no/maybe reasoning over biomedical abstracts
MMLU-Med
92.5
+5.5vs MedGemmaMedical subset of MMLU knowledge tasks
MedXpertQA
41.0
+15.3vs MedGemmaExpert-level clinical reasoning, 17 specialties · text
Deltas are against MedGemma 27B’s published figures. Where Nia sits against the wider field is shown benchmark by benchmark below.
Macro average across all five text benchmarks
Nia against the medical peer
MedGemma 27B is the reference point that matters most: a medical-specialised model at essentially the same size. Its published figures use test-time scaling by its own model card’s admission, while every Nia figure here is a single greedy pass — so this runs against its best-case settings, not ours.
- Nia 1.0
- MedGemma 27B — published
MedQA
MedMCQA
PubMedQA
MMLU-Med
MedXpertQA
View as table
| Benchmark | Nia 1.0 | MedGemma 27B — published |
|---|---|---|
| MedQA | 93.0 | 87.7 |
| MedMCQA | 76.5 | 74.2 |
| PubMedQA | 71.5 | 76.8 |
| MMLU-Med | 92.5 | 87.0 |
| MedXpertQA | 41.0 | 25.7 |
Macro average across all five
MedXpertQA
accuracyExpert-level clinical reasoning · text
41.0+3.2DeepSeek-R1671B MoE37.8The best published MedXpertQA text score — from a model roughly 24× larger.
MedQA
accuracyAgainst the UAE incumbent
93.0+5.7Med42-v2 70B70B87.3M42's clinical model, on USMLE sample questions — at 2.5× the parameters.
MMLU-Med
accuracyMedical knowledge subset
92.5+2.5Gemini 2.5 Prosize not published90.0Ahead of Google's frontier model on its published medical figure.
VQA-RAD
tokenised F1Radiology visual question answering
66.8+20.1MedGemma 27B27B46.7+20.1 over the published figure, at a comparable model size.
Benchmark by benchmark
Benchmark by benchmark, against the medical field
One panel per benchmark. Each shows every medical model that reports it — MedGemma, Med42-v2, Med-PaLM 2, MedGemma 1.5 — and then the general-purpose frontier models Nia is ahead of on that benchmark. Nia leads the medical field on seven of the eight.
- #1 of 5
MedQA
USMLE-style medical licensing exam questions
- Nia 1.0R93.0
- MedGemma 27B87.7
- Med42-v2 70B87.3†
- Med-PaLM 286.5
- MedGemma 1.5 4B69.1
General models Nia leads here
- GPT-5.4R92.1
- DeepSeek-R1R89.8
- DeepSeek-V3.189.5
- Claude Opus 4.587.4
- DeepSeek R2R~85.0
- Qwen 3.5~82.0
- Llama 4 Maverick~79.0
accuracy · Nia measured, others published. Medical models in full; general models shown where Nia is ahead — the complete field, in order, is in the table below.
- #1 of 3
MedMCQA
Indian medical entrance exam, multi-option
- Nia 1.0R76.5
- MedGemma 27B74.2
- MedGemma 1.5 4B59.8
General models Nia leads here
- Qwen 3.5~76.0
- Llama 4 Maverick~72.0
accuracy · Nia measured, others published. Medical models in full; general models shown where Nia is ahead — the complete field, in order, is in the table below.
- #3 of 4
PubMedQA
Yes/no/maybe reasoning over biomedical abstracts
- Med-PaLM 279.0
- MedGemma 27B76.8
- Nia 1.0R71.5
- MedGemma 1.5 4B68.2
General models Nia leads here
- Llama 4 Maverick~70.0
accuracy · Nia measured, others published. Medical models in full; general models shown where Nia is ahead — the complete field, in order, is in the table below.
- #1 of 3
MMLU-Med
Medical subset of MMLU knowledge tasks
- Nia 1.0R92.5
- MedGemma 27B87.0
- MedGemma 1.5 4B69.6
General models Nia leads here
- Gemini 2.5 ProR~90.0
- DeepSeek R2R~85.0
- Qwen 3.5~83.0
- Llama 4 Maverick~79.0
accuracy · Nia measured, others published. Medical models in full; general models shown where Nia is ahead — the complete field, in order, is in the table below.
- #1 of 3
MedXpertQA
Expert-level clinical reasoning, 17 specialties · text
- Nia 1.0R41.0
- MedGemma 27B25.7
- MedGemma 1.5 4B16.4
General models Nia leads here
- DeepSeek-R1R37.8
- o3-miniR37.3
accuracy · Nia measured, others published. Medical models in full; general models shown where Nia is ahead — the complete field, in order, is in the table below.
Panels are derived from the same data as the table below. The medical field is shown in full; general-purpose models appear where Nia is ahead of them, and the complete field in true order — including the models that lead Nia — is in the table further down. Bars are scaled within each panel and never read across. Nia is measured; every other figure is published, and rows marked ~ are approximate by their own source’s label.
The frontier
Frontier medical accuracy, at a fraction of the size
Plotting MedQA accuracy against parameter count shows what a table can only imply: among every model that publishes its size, nothing beats Nia — including 671B mixture-of-experts systems roughly 24× larger.
The frontier ends at Nia. The accent line traces the best MedQA accuracy reachable at each parameter count. Every larger model plotted here — up to 671B mixture-of-experts — sits below it.
Only models whose parameter count is published can be plotted. DeepSeek-R1 (89.8) and V3.1 (89.5) are both 671B, so their marks overlap and their labels are offset with leader lines.
Full competitive context
Healthcare models and current-generation frontier flagships, ordered by MedQA. Nia is not the top row on every column, and the table shows that plainly.
| Model | Params | MedQA | MedMCQA | PubMedQA | MMLU-Med | MedXpert | Source |
|---|---|---|---|---|---|---|---|
| Nia 1.0Infinia TechnologiesR | 28B | 93.0 | 76.5 | 71.5 | 92.5 | 41.0 | measured, this release |
| MedGemma 27BGoogle | 27B | 87.7 | 74.2 | 76.8 | 87.0 | 25.7 | MedGemma model card (Google) |
| Med42-v2 70BM42 · Abu Dhabi | 70B | 87.3† | — | — | — | — | Med42-v2 paper (M42) |
| Med-PaLM 2Google | 340B | 86.5 | — | 79.0 | — | — | Med-PaLM 2 paper (Google) |
| MedGemma 1.5 4BGoogle | 4B | 69.1 | 59.8 | 68.2 | 69.6 | 16.4 | MedGemma 1.5 technical report (Google) |
| GPT-5OpenAIR | — | ~93.0 | ~87.0 | ~80.0 | ~93.0 | — | Medical LLM Leaderboard |
| GPT-5.4OpenAIR | — | 92.1 | 79.1 | — | — | — | MedQA leaderboard |
| Gemini 2.5 ProGoogleR | — | ~94.6 | ~83.0 | ~80.0 | ~90.0 | — | MedQA leaderboard |
| Claude Opus 4.5Anthropic | — | 87.4 | 78.7 | — | — | — | MedQA leaderboard |
| DeepSeek-V3.1DeepSeek | 671B MoE | 89.5 | 79.0 | — | — | — | MedQA leaderboard |
| DeepSeek R2DeepSeekR | — | ~85.0 | ~78.0 | ~74.0 | ~85.0 | — | Medical LLM Leaderboard |
| Qwen 3.5Alibaba (Qwen) | — | ~82.0 | ~76.0 | ~72.0 | ~83.0 | — | Medical LLM Leaderboard |
| Llama 4 MaverickMeta | 400B MoE | ~79.0 | ~72.0 | ~70.0 | ~79.0 | — | Medical LLM Leaderboard |
| DeepSeek-R1DeepSeekR | 671B MoE | 89.8 | 78.4 | — | — | 37.8 | MedQA leaderboard |
| o3-miniOpenAIR | — | — | — | — | — | 37.3 | official MedXpertQA leaderboard |
Healthcare models and current-generation frontier flagships. Only Nia’s row is measured by us; every other figure is that vendor’s, paper’s or leaderboard’s published number, gathered under protocols that differ and are mostly unstated. † marks a figure from a related but different test — Med42-v2’s is the USMLE sample exam, not MedQA proper. Treat this as indicative context rather than a controlled measurement.
Multimodal
It reads images as well as text
Nia takes radiology images alongside text and answers questions about them. Image understanding was not separately trained, so these results reflect the model's inherited multimodal ability — and they still lead MedGemma 27B on all three imaging benchmarks, as well as Google's newest medical release, MedGemma 1.5.
- Nia 1.0
- MedGemma 27B — published
MedXpertQA-MM
accuracy
SLAKE
tokenised F1
VQA-RAD
tokenised F1
View as table
| Benchmark | Nia 1.0 | MedGemma 27B — published |
|---|---|---|
| MedXpertQA-MM (accuracy) | 40.0 | 26.8 |
| SLAKE (tokenised F1) | 72.3 | 70.0 |
| VQA-RAD (tokenised F1) | 66.8 | 46.7 |
MedXpertQA-MM is accuracy; SLAKE and VQA-RAD are tokenised F1. Different scales, shown on one axis for compactness and labelled per row — never compared across rows.
- #1 of 3
MedXpertQA-MM
Expert multimodal clinical reasoning
- Nia 1.040.0
- MedGemma 27B26.8
- MedGemma 1.5 4B20.9
General models Nia leads here
- Gemini-2.0-Flash37.2
accuracy · Nia measured, others published. Medical models in full; general models shown where Nia is ahead — the complete field, in order, is in the table below.
- #1 of 3
SLAKE
Semantically-labelled radiology VQA
- Nia 1.072.3
- MedGemma 27B70.0
- MedGemma 1.5 4B59.7
tokenised F1 · Nia measured, others published. Medical models in full; general models shown where Nia is ahead — the complete field, in order, is in the table below.
- #1 of 3
VQA-RAD
Clinician-generated radiology VQA
- Nia 1.066.8
- MedGemma 1.5 4B48.1
- MedGemma 27B46.7
tokenised F1 · Nia measured, others published. Medical models in full; general models shown where Nia is ahead — the complete field, in order, is in the table below.
| Model | MedXpertQA-MMaccuracy | SLAKEtok. F1 | VQA-RADtok. F1 | Source |
|---|---|---|---|---|
| Nia 1.0Infinia Technologies | 40.0 | 72.3 | 66.8 | measured, this release |
| MedGemma 27BGoogle | 26.8 | 70.0 | 46.7 | MedGemma model card (Google) |
| MedGemma 1.5 4BGoogle | 20.9 | 59.7 | 48.1 | MedGemma 1.5 technical report (Google) |
| o1OpenAIR | 56.3 | — | — | official MedXpertQA leaderboard |
| GPT-4oOpenAI | 42.8 | — | — | official MedXpertQA leaderboard |
| Gemini-2.0-FlashGoogle | 37.2 | — | — | official MedXpertQA leaderboard |
SLAKE and VQA-RAD are tokenised F1; MedXpertQA-MM is accuracy. These are not the same scale and are never compared across columns. Nia’s image understanding was not separately trained, so these figures reflect inherited multimodal ability.
Why it leads
What the numbers rest on
Clears the medical peer on 4 of 5, by up to +15.3
MedGemma 27B is a medical-specialised model at essentially the same size, which makes it the reference point that matters. Nia is ahead of its published figures on 4 of the 5 text benchmarks, with a macro average of 74.9 against 70.3. PubMedQA is the exception, and it is shown.
Up to 24× smaller than the models it matches
The systems that trade places with Nia at the top are 400–671B mixture-of-experts models, or frontier models whose vendors publish no size at all. Nia is a 28B dense model that runs on a single box — the accuracy-per-parameter gap is the whole point.
No test-time scaling, no retrieval, no ensembling
Every Nia figure is one greedy pass at temperature 0, zero-shot. MedGemma's published numbers use test-time scaling by its own card's admission — it reports MedQA as 89.8 best-of-5 versus 87.7 zero-shot. Like-for-like, Nia is being compared against others' best-case settings.
Decontaminated against every evaluation split
Training data was screened against all eval splits with 13-gram overlap plus exact and short-containment matching, and benchmark splits were held out entirely. Contamination inflates published medical numbers industry-wide; we can at least show what we did about ours.
The flagship
Nia by Infinia
Infinia Technologies' flagship medical model — built in Abu Dhabi by New Emerging Technologies in collaboration with Apeiro Digital, the healthcare and medical arm of Infinia Technologies. One group, accountable from training to deployment.
Built in the United Arab Emirates
Sovereign by design
Trained in Abu Dhabi
Nia was trained end to end by New Emerging Technologies in Abu Dhabi. Data handling, evaluation and decontamination sit with one accountable team in the UAE — provenance you can audit, not a vendor's black box.
Your infrastructure, not someone else's API
Weights are delivered to qualified organisations under an agreement. Nia runs inside your own perimeter — no clinical workload has to depend on an overseas endpoint.
Data residency by architecture
A 28B dense model runs on a single box, so a hospital, insurer or government can deploy Nia entirely within its own infrastructure. Patient data stays where the law — and the patient — expect it to.
Access
Evaluate it, or put it to work
Whether you want to try Nia, run it through your own evaluation, or deploy it against a real use case, the route starts the same way. Weights go to qualified organisations under an agreement; implementation runs through Apeiro Digital. Tell us what you need and we'll be in touch from Abu Dhabi.
- Full model weights in bfloat16, text and image input — run entirely on your own infrastructure.
- The complete model card, including the benchmark we lose.
- Our evaluation setup, so you can reproduce every figure here.
- Deployment, integration and support with Apeiro Digital when you move to production.
Deployment & integration
Implementation and integration are delivered in collaboration with Apeiro Digital, the healthcare and medical arm of Infinia Technologies — so evaluation, deployment and support sit with one group.