Abu Dhabi · United Arab EmiratesReleased August 2026

Nia 1.0

Frontier medical intelligence, built in the UAE.

A 28-billion-parameter multimodal healthcare language model — clinical reasoning, medical examinations, literature and imaging, from a model small enough to run on a single machine, entirely inside your own infrastructure.

parameters
28Bparameters
multimodal
Text + imagemultimodal
context
262Kcontext

Weights are available to qualified organisations under an evaluation agreement. Nia is a clinical decision support tool — not a licensed clinician, and not a substitute for professional medical judgement.

MedQAUSMLE-style · accuracy
93.0

Ahead of every medical model with a published figure — including Google’s MedGemma and M42’s Med42.

  • GoogleMedGemma 27BGoogle87.7
  • M42 · Abu DhabiMed42-v2 70BM42 · Abu Dhabi87.3
  • GoogleMed-PaLM 2Google86.5
MedXpertQA · text41.0

Ahead of the best published text results — DeepSeek-R1 37.8 and o3-mini 37.3, per the official leaderboard.

Nia measured in bfloat16, single greedy pass, temperature 0, zero-shot. Competitor figures are published by their vendors.

Benchmarked against

  • GoogleMedGemma 27B
  • M42Med42-v2 70B
  • GoogleMed-PaLM 2
  • GoogleMedGemma 1.5 4B
  • OpenAIGPT-5
  • OpenAIGPT-5.4
  • GoogleGemini 2.5 Pro
  • AnthropicClaude Opus 4.5
  • DeepSeekDeepSeek-V3.1
  • DeepSeekDeepSeek R2
  • Alibaba (Qwen)Qwen 3.5
  • MetaLlama 4 Maverick
  • DeepSeekDeepSeek-R1
  • OpenAIo3-mini
  • OpenAIo1
  • OpenAIGPT-4o
  • GoogleGemini-2.0-Flash

Benchmarks

Measured against the models built for medicine

The comparison that matters most is against other healthcare models. Nia clears MedGemma 27B’s published figures on 4 of 5 text benchmarks — by up to +15.3 points — and is ahead of Google’s Med-PaLM 2 and M42’s Med42-v2 on MedQA. PubMedQA is the one it loses, and it is on this page for that reason.

  • MedQA

    93.0

    +5.3vs MedGemma

    USMLE-style medical licensing exam questions

  • MedMCQA

    76.5

    +2.3vs MedGemma

    Indian medical entrance exam, multi-option

  • PubMedQA

    71.5

    5.3vs MedGemma

    Yes/no/maybe reasoning over biomedical abstracts

  • MMLU-Med

    92.5

    +5.5vs MedGemma

    Medical subset of MMLU knowledge tasks

  • MedXpertQA

    41.0

    +15.3vs MedGemma

    Expert-level clinical reasoning, 17 specialties · text

Deltas are against MedGemma 27B’s published figures. Where Nia sits against the wider field is shown benchmark by benchmark below.

Macro average across all five text benchmarks

MedGemma 27B 70.374.9+4.6

Nia against the medical peer

MedGemma 27B is the reference point that matters most: a medical-specialised model at essentially the same size. Its published figures use test-time scaling by its own model card’s admission, while every Nia figure here is a single greedy pass — so this runs against its best-case settings, not ours.

  • Nia 1.0
  • MedGemma 27B — published
  • MedQA

    93.0
    87.7
  • MedMCQA

    76.5
    74.2
  • PubMedQA

    71.5
    76.8
  • MMLU-Med

    92.5
    87.0
  • MedXpertQA

    41.0
    25.7
View as table
Text benchmark accuracy for Nia 1.0 against MedGemma 27B's published figures.
BenchmarkNia 1.0MedGemma 27B — published
MedQA93.087.7
MedMCQA76.574.2
PubMedQA71.576.8
MMLU-Med92.587.0
MedXpertQA41.025.7

Macro average across all five

Nia 74.9MedGemma 27B 70.3
  • MedXpertQA

    accuracy

    Expert-level clinical reasoning · text

    41.0+3.2
    DeepSeekDeepSeek-R1671B MoE37.8

    The best published MedXpertQA text score — from a model roughly 24× larger.

  • MedQA

    accuracy

    Against the UAE incumbent

    93.0+5.7
    M42 · Abu DhabiMed42-v2 70B70B87.3

    M42's clinical model, on USMLE sample questions — at 2.5× the parameters.

  • MMLU-Med

    accuracy

    Medical knowledge subset

    92.5+2.5
    GoogleGemini 2.5 Prosize not published90.0

    Ahead of Google's frontier model on its published medical figure.

  • VQA-RAD

    tokenised F1

    Radiology visual question answering

    66.8+20.1
    GoogleMedGemma 27B27B46.7

    +20.1 over the published figure, at a comparable model size.

Benchmark by benchmark

Benchmark by benchmark, against the medical field

One panel per benchmark. Each shows every medical model that reports it — MedGemma, Med42-v2, Med-PaLM 2, MedGemma 1.5 — and then the general-purpose frontier models Nia is ahead of on that benchmark. Nia leads the medical field on seven of the eight.

  • MedQA

    USMLE-style medical licensing exam questions

    #1 of 5
    1. Nia 1.0R93.0
    2. GoogleMedGemma 27B87.7
    3. M42 · Abu DhabiMed42-v2 70B87.3
    4. GoogleMed-PaLM 286.5
    5. GoogleMedGemma 1.5 4B69.1

    General models Nia leads here

    1. OpenAIGPT-5.4R92.1
    2. DeepSeekDeepSeek-R1R89.8
    3. DeepSeekDeepSeek-V3.189.5
    4. AnthropicClaude Opus 4.587.4
    5. DeepSeekDeepSeek R2R~85.0
    6. Alibaba (Qwen)Qwen 3.5~82.0
    7. MetaLlama 4 Maverick~79.0

    accuracy · Nia measured, others published. Medical models in full; general models shown where Nia is ahead — the complete field, in order, is in the table below.

  • MedMCQA

    Indian medical entrance exam, multi-option

    #1 of 3
    1. Nia 1.0R76.5
    2. GoogleMedGemma 27B74.2
    3. GoogleMedGemma 1.5 4B59.8

    General models Nia leads here

    1. Alibaba (Qwen)Qwen 3.5~76.0
    2. MetaLlama 4 Maverick~72.0

    accuracy · Nia measured, others published. Medical models in full; general models shown where Nia is ahead — the complete field, in order, is in the table below.

  • PubMedQA

    Yes/no/maybe reasoning over biomedical abstracts

    #3 of 4
    1. GoogleMed-PaLM 279.0
    2. GoogleMedGemma 27B76.8
    3. Nia 1.0R71.5
    4. GoogleMedGemma 1.5 4B68.2

    General models Nia leads here

    1. MetaLlama 4 Maverick~70.0

    accuracy · Nia measured, others published. Medical models in full; general models shown where Nia is ahead — the complete field, in order, is in the table below.

  • MMLU-Med

    Medical subset of MMLU knowledge tasks

    #1 of 3
    1. Nia 1.0R92.5
    2. GoogleMedGemma 27B87.0
    3. GoogleMedGemma 1.5 4B69.6

    General models Nia leads here

    1. GoogleGemini 2.5 ProR~90.0
    2. DeepSeekDeepSeek R2R~85.0
    3. Alibaba (Qwen)Qwen 3.5~83.0
    4. MetaLlama 4 Maverick~79.0

    accuracy · Nia measured, others published. Medical models in full; general models shown where Nia is ahead — the complete field, in order, is in the table below.

  • MedXpertQA

    Expert-level clinical reasoning, 17 specialties · text

    #1 of 3
    1. Nia 1.0R41.0
    2. GoogleMedGemma 27B25.7
    3. GoogleMedGemma 1.5 4B16.4

    General models Nia leads here

    1. DeepSeekDeepSeek-R1R37.8
    2. OpenAIo3-miniR37.3

    accuracy · Nia measured, others published. Medical models in full; general models shown where Nia is ahead — the complete field, in order, is in the table below.

Panels are derived from the same data as the table below. The medical field is shown in full; general-purpose models appear where Nia is ahead of them, and the complete field in true order — including the models that lead Nia — is in the table further down. Bars are scaled within each panel and never read across. Nia is measured; every other figure is published, and rows marked ~ are approximate by their own source’s label.

The frontier

Frontier medical accuracy, at a fraction of the size

Plotting MedQA accuracy against parameter count shows what a table can only imply: among every model that publishes its size, nothing beats Nia — including 671B mixture-of-experts systems roughly 24× larger.

M4Nia 1.0 93.0DeepSeek-R1 89.8DeepSeek-V3.1 89.5MedGemma 27B 87.7Med42-v2 70B 87.3Med-PaLM 2 86.5Llama 4 Maverick 79.0MedGemma 1.5 4B 69.1

The frontier ends at Nia. The accent line traces the best MedQA accuracy reachable at each parameter count. Every larger model plotted here — up to 671B mixture-of-experts — sits below it.

Only models whose parameter count is published can be plotted. DeepSeek-R1 (89.8) and V3.1 (89.5) are both 671B, so their marks overlap and their labels are offset with leader lines.

Full competitive context

Healthcare models and current-generation frontier flagships, ordered by MedQA. Nia is not the top row on every column, and the table shows that plainly.

Published medical benchmark figures for frontier and open models, compared with Nia 1.0’s measured results. Figures marked with a tilde are approximate by their source’s own label; an R badge marks reasoning models.
ModelParamsMedQAMedMCQAPubMedQAMMLU-MedMedXpertSource
Infinia TechnologiesNia 1.0Infinia TechnologiesR28B93.076.571.592.541.0measured, this release
GoogleMedGemma 27BGoogle27B87.774.276.887.025.7MedGemma model card (Google)
M42 · Abu DhabiMed42-v2 70BM42 · Abu Dhabi70B87.3Med42-v2 paper (M42)
GoogleMed-PaLM 2Google340B86.579.0Med-PaLM 2 paper (Google)
GoogleMedGemma 1.5 4BGoogle4B69.159.868.269.616.4MedGemma 1.5 technical report (Google)
OpenAIGPT-5OpenAIR~93.0~87.0~80.0~93.0Medical LLM Leaderboard
OpenAIGPT-5.4OpenAIR92.179.1MedQA leaderboard
GoogleGemini 2.5 ProGoogleR~94.6~83.0~80.0~90.0MedQA leaderboard
AnthropicClaude Opus 4.5Anthropic87.478.7MedQA leaderboard
DeepSeekDeepSeek-V3.1DeepSeek671B MoE89.579.0MedQA leaderboard
DeepSeekDeepSeek R2DeepSeekR~85.0~78.0~74.0~85.0Medical LLM Leaderboard
Alibaba (Qwen)Qwen 3.5Alibaba (Qwen)~82.0~76.0~72.0~83.0Medical LLM Leaderboard
MetaLlama 4 MaverickMeta400B MoE~79.0~72.0~70.0~79.0Medical LLM Leaderboard
DeepSeekDeepSeek-R1DeepSeekR671B MoE89.878.437.8MedQA leaderboard
OpenAIo3-miniOpenAIR37.3official MedXpertQA leaderboard

Healthcare models and current-generation frontier flagships. Only Nia’s row is measured by us; every other figure is that vendor’s, paper’s or leaderboard’s published number, gathered under protocols that differ and are mostly unstated. marks a figure from a related but different test — Med42-v2’s is the USMLE sample exam, not MedQA proper. Treat this as indicative context rather than a controlled measurement.

Multimodal

It reads images as well as text

Nia takes radiology images alongside text and answers questions about them. Image understanding was not separately trained, so these results reflect the model's inherited multimodal ability — and they still lead MedGemma 27B on all three imaging benchmarks, as well as Google's newest medical release, MedGemma 1.5.

  • Nia 1.0
  • MedGemma 27B — published
  • MedXpertQA-MM

    accuracy

    40.0
    26.8
  • SLAKE

    tokenised F1

    72.3
    70.0
  • VQA-RAD

    tokenised F1

    66.8
    46.7
View as table
Imaging benchmark results for Nia 1.0 against MedGemma 27B's published figures. Metrics differ per benchmark.
BenchmarkNia 1.0MedGemma 27B — published
MedXpertQA-MM (accuracy)40.026.8
SLAKE (tokenised F1)72.370.0
VQA-RAD (tokenised F1)66.846.7

MedXpertQA-MM is accuracy; SLAKE and VQA-RAD are tokenised F1. Different scales, shown on one axis for compactness and labelled per row — never compared across rows.

  • MedXpertQA-MM

    Expert multimodal clinical reasoning

    #1 of 3
    1. Nia 1.040.0
    2. GoogleMedGemma 27B26.8
    3. GoogleMedGemma 1.5 4B20.9

    General models Nia leads here

    1. GoogleGemini-2.0-Flash37.2

    accuracy · Nia measured, others published. Medical models in full; general models shown where Nia is ahead — the complete field, in order, is in the table below.

  • SLAKE

    Semantically-labelled radiology VQA

    #1 of 3
    1. Nia 1.072.3
    2. GoogleMedGemma 27B70.0
    3. GoogleMedGemma 1.5 4B59.7

    tokenised F1 · Nia measured, others published. Medical models in full; general models shown where Nia is ahead — the complete field, in order, is in the table below.

  • VQA-RAD

    Clinician-generated radiology VQA

    #1 of 3
    1. Nia 1.066.8
    2. GoogleMedGemma 1.5 4B48.1
    3. GoogleMedGemma 27B46.7

    tokenised F1 · Nia measured, others published. Medical models in full; general models shown where Nia is ahead — the complete field, in order, is in the table below.

MedXpertQA-MM, SLAKE and VQA-RAD figures for multimodal models, compared with Nia 1.0.
ModelMedXpertQA-MMaccuracySLAKEtok. F1VQA-RADtok. F1Source
Infinia TechnologiesNia 1.0Infinia Technologies40.072.366.8measured, this release
GoogleMedGemma 27BGoogle26.870.046.7MedGemma model card (Google)
GoogleMedGemma 1.5 4BGoogle20.959.748.1MedGemma 1.5 technical report (Google)
OpenAIo1OpenAIR56.3official MedXpertQA leaderboard
OpenAIGPT-4oOpenAI42.8official MedXpertQA leaderboard
GoogleGemini-2.0-FlashGoogle37.2official MedXpertQA leaderboard

SLAKE and VQA-RAD are tokenised F1; MedXpertQA-MM is accuracy. These are not the same scale and are never compared across columns. Nia’s image understanding was not separately trained, so these figures reflect inherited multimodal ability.

Why it leads

What the numbers rest on

  • Clears the medical peer on 4 of 5, by up to +15.3

    MedGemma 27B is a medical-specialised model at essentially the same size, which makes it the reference point that matters. Nia is ahead of its published figures on 4 of the 5 text benchmarks, with a macro average of 74.9 against 70.3. PubMedQA is the exception, and it is shown.

  • Up to 24× smaller than the models it matches

    The systems that trade places with Nia at the top are 400–671B mixture-of-experts models, or frontier models whose vendors publish no size at all. Nia is a 28B dense model that runs on a single box — the accuracy-per-parameter gap is the whole point.

  • No test-time scaling, no retrieval, no ensembling

    Every Nia figure is one greedy pass at temperature 0, zero-shot. MedGemma's published numbers use test-time scaling by its own card's admission — it reports MedQA as 89.8 best-of-5 versus 87.7 zero-shot. Like-for-like, Nia is being compared against others' best-case settings.

  • Decontaminated against every evaluation split

    Training data was screened against all eval splits with 13-gram overlap plus exact and short-containment matching, and benchmark splits were held out entirely. Contamination inflates published medical numbers industry-wide; we can at least show what we did about ours.

The flagship

Nia by Infinia

Infinia Technologies' flagship medical model — built in Abu Dhabi by New Emerging Technologies in collaboration with Apeiro Digital, the healthcare and medical arm of Infinia Technologies. One group, accountable from training to deployment.

Built in the United Arab Emirates

Sovereign by design

  • Trained in Abu Dhabi

    Nia was trained end to end by New Emerging Technologies in Abu Dhabi. Data handling, evaluation and decontamination sit with one accountable team in the UAE — provenance you can audit, not a vendor's black box.

  • Your infrastructure, not someone else's API

    Weights are delivered to qualified organisations under an agreement. Nia runs inside your own perimeter — no clinical workload has to depend on an overseas endpoint.

  • Data residency by architecture

    A 28B dense model runs on a single box, so a hospital, insurer or government can deploy Nia entirely within its own infrastructure. Patient data stays where the law — and the patient — expect it to.

Access

Evaluate it, or put it to work

Whether you want to try Nia, run it through your own evaluation, or deploy it against a real use case, the route starts the same way. Weights go to qualified organisations under an agreement; implementation runs through Apeiro Digital. Tell us what you need and we'll be in touch from Abu Dhabi.

  • Full model weights in bfloat16, text and image input — run entirely on your own infrastructure.
  • The complete model card, including the benchmark we lose.
  • Our evaluation setup, so you can reproduce every figure here.
  • Deployment, integration and support with Apeiro Digital when you move to production.

Deployment & integration

Implementation and integration are delivered in collaboration with Apeiro Digital, the healthcare and medical arm of Infinia Technologies — so evaluation, deployment and support sit with one group.

We use these details only to assess your request.