OMNI
Scaling Multi-Teacher Distillation for Digital Pathology
1Université Paris-Saclay, CentraleSupélec, Gustave Roussy, INSERM, IHU PRISM, Cancer Data Science Unit, France
2Université Paris-Saclay, CentraleSupélec, MICS Laboratory, France
3Stony Brook University, USA
*Equal contribution
NeurIPS 2026, Main Track
Abstract
Multi-teacher distillation allows transferring knowledge from multiple “teacher” networks to a single “student” network. This method is promising in fields such as digital pathology, where many powerful foundation models were proposed recently. However, as we show in this paper, the standard multi-teacher distillation approach reacts poorly to an increase in the number of teachers used to train a student encoder. We demonstrate that learning teacher-specific representations is a key to scaling in MTD. Importantly, different design choices, such as learnable teacher tokens, a tailored attention scheme, additional mixture-of-experts layers and a contrastive loss, are proposed to better learn such teacher-specific representations. We show that our method better scales with respect to the number of considered teachers, allowing us to train compact student encoders, up to 10× more efficient than larger teacher foundation models, while matching their performance and even outperforming them on a large set of tile-level and slide-level tasks. We conduct a thorough empirical validation, evaluating more than 10 foundation models on 39 tasks spanning tile and slide levels. As a result, we release a new collection of strong and efficient foundation models, named OMNI, trained from the knowledge of 10 state-of-the-art foundation models. We finally conduct an analysis of the learned teacher-specific representations, highlighting their complementarity and explaining why they can be easily aggregated at downstream time.
Vanilla MTD
All teachers are reconstructed from a single shared CLS token.
OMNI (ours)
Each teacher is reconstructed from its own teacher-specific token.
Method
Vanilla multi-teacher distillation
Following previous work, a ViT student
where losses combine a cosine and a smooth L1 term, as done in previous work:
A single shared CLS token must thus simultaneously compress all teachers into one representation, which becomes increasingly bottlenecked as
Teacher-specific tokens
We therefore abandon the single-token design and instead assign one independent token per teacher. Formally, we maintain a learnable token matrix
Structured attention masking
To encourage the learning of a specific aggregation of patch token embeddings for each teacher, we apply the following block-structured attention mask in all layers. Each teacher token only attends to itself among teacher tokens, which guarantees that it is exclusively informed by the patch tokens rather than by the already-aggregated views of other teachers.
MoE blocks and training objective
As a complementary mechanism, the last transformer blocks (3 for OMNI-T/S/B, 5 for OMNI-L) replace the standard dense feed-forward networks with Mixture-of-Experts (MoE) layers, with 5 experts and top-2 routing. This increases the total model capacity while keeping inference cost constant, and a Switch Transformer load-balancing loss encourages balanced expert utilization. We also add a symmetric InfoNCE loss for each teacher, penalizing embeddings that collapse to similar representations across different samples. The full training objective is
with
Results
Scaling with the number of teachers
For a fair evaluation, we rank teachers by their individual performance on the THUNDER benchmark and add them from strongest to weakest. Increasing the number of teachers for vanilla MTD does not improve downstream performance, and even degrades it, while our method better scales. Similarly, the reconstruction quality of UNI2-h drops from 0.90 to 0.83 cosine similarity for vanilla MTD, whereas our approach retains 0.87 at 10 teachers.
Linear probing benchmarking
The full OMNI family demonstrates a consistent scaling trend. OMNI-B and OMNI-L are ranked second and first respectively, while being up to 10× more efficient in GFLOPs than the best teacher models. Notably, OMNI-L achieves the best tile-level F1 (87.8%) and slide-level BAcc (63.5%), outperforming all teachers including UNI2-h (681M) and H-optimus-1 (1.1B) despite being significantly smaller. A paired Wilcoxon signed-rank test shows that OMNI-L significantly outperforms the large majority of teachers at both tile and slide level.
| Model | Params | GFLOPs | Tile F1 (%) 13 datasets | Tile rank ↓ | Slide BAcc (%) 25 tasks | Slide rank ↓ | Mean score | Mean rank ↓ |
|---|---|---|---|---|---|---|---|---|
| Foundation models | ||||||||
| UNI2-h | 681M | 180.4 | 87.4 | 4.6 | 61.5 | 9.0 | 74.5 | 6.8 |
| Virchow2 | 631M | 164.6 | 85.9 | 6.6 | 61.5 | 8.3 | 73.7 | 7.4 |
| H-optimus-1 | 1.1B | 296.0 | 86.3 | 6.2 | 62.5 | 4.7 | 74.4 | 5.4 |
| KEEP | 414M | 59.7 | 84.9 | 9.5 | 60.7 | 11.0 | 72.8 | 10.2 |
| Prov-GigaPath | 1.1B | 223.4 | 84.2 | 9.3 | 61.0 | 10.7 | 72.6 | 10.0 |
| CONCH 1.5 | 307M | 79.5 | 84.3 | 11.0 | 61.5 | 9.4 | 72.9 | 10.2 |
| Hibou-B | 86M | 22.0 | 82.5 | 10.7 | 60.3 | 11.0 | 71.4 | 10.8 |
| Kaiko-ViT-B/8 | 86M | 66.9 | 82.0 | 13.2 | 58.9 | 13.9 | 70.5 | 13.6 |
| Distilled models | ||||||||
| GPFM | 303M | 81.1 | 83.2 | 9.5 | 60.9 | 9.1 | 72.0 | 9.3 |
| H0-mini | 86M | 22.3 | 84.9 | 8.3 | 61.5 | 9.4 | 73.2 | 8.9 |
| Vanilla MTD, ViT-B | 86M | 22.0 | 85.4 | 5.9 | 61.8 | 8.7 | 73.6 | 7.3 |
| Ours | ||||||||
| OMNI-T | 7.3M | 1.4 | 83.9 | 11.0 | 63.2 | 4.0 | 73.6 | 7.5 |
| OMNI-S | 28.7M | 5.5 | 85.3 | 7.6 | 62.9 | 4.1 | 74.1 | 5.8 |
| OMNI-B | 114M | 21.8 | 87.1 | 4.0 | 63.2 | 4.4 | 75.2 | 4.2 |
| OMNI-L | 330M | 67.5 | 87.8 | 2.5 | 63.5 | 2.1 | 75.7 | 2.3 |
Linear probing performance at tile level (F1, 13 datasets) and slide level (BAcc, 7 CPTAC datasets / 25 tasks), and their mean. Mean rank is computed per dataset then averaged. Bold: best; underline: second best. TCGA-derived THUNDER datasets are removed to avoid data leakage.
Additional results: KNN and 16-shot classification
| Model | KNN F1 (%) | KNN rank ↓ | 16-shot F1 (%) | 16-shot rank ↓ |
|---|---|---|---|---|
| UNI2-h | 85.0 | 4.7 | 81.9 | 3.4 |
| Virchow2 | 84.7 | 5.0 | 75.9 | 9.6 |
| H-optimus-1 | 83.9 | 5.3 | 78.3 | 7.1 |
| KEEP | 83.4 | 7.7 | 79.3 | 7.4 |
| Prov-GigaPath | 81.3 | 9.1 | 77.2 | 8.3 |
| CONCH 1.5 | 82.0 | 9.4 | 77.2 | 9.1 |
| Kaiko-ViT-B/8 | 80.0 | 10.5 | 78.6 | 7.5 |
| GPFM | 82.2 | 7.9 | 77.8 | 7.2 |
| H0-mini | 81.4 | 8.4 | 76.6 | 8.8 |
| OMNI-T | 81.5 | 8.8 | 77.5 | 8.1 |
| OMNI-S | 82.3 | 6.9 | 78.7 | 6.0 |
| OMNI-B | 84.3 | 3.6 | 79.9 | 4.8 |
| OMNI-L | 84.8 | 3.6 | 80.7 | 3.7 |
Ablation studies
| Variant of OMNI-B | InfoNCE | MoE tail | Attention mask | Mean F1 (%) |
|---|---|---|---|---|
| OMNI-B (full) | ✓ | ✓ | ✓ | 87.1 |
| No InfoNCE loss | ✗ | ✓ | ✓ | 86.6−0.5 |
| No MoE tail | ✓ | ✗ | ✓ | 86.6−0.5 |
| No attention mask | ✓ | ✓ | ✗ | 86.2−0.9 |
| Vanilla MTD | ✗ | ✗ | ✗ | 85.4−1.7 |
Ablation study on the contribution of the structured attention, InfoNCE loss and MoE blocks (mean F1 on THUNDER linear probing, ViT-Base). When removing each contribution, performance drops, showing that all of them are necessary.
Uncertainty and robustness
| Model | F1 drop under PGD ↓ | ECE ↓ |
|---|---|---|
| UNI2-h | 30.4 | 4.0 |
| Virchow2 | 29.4 | 3.8 |
| H-optimus-1 | 58.1 | 4.0 |
| KEEP | 43.9 | 4.4 |
| H0-mini | 33.9 | 4.3 |
| Kaiko-ViT-B/8 | 39.4 | 3.5 |
| Prov-GigaPath | 40.8 | 3.7 |
| CONCH 1.5 | 77.1 | 4.7 |
| OMNI-T | 31.6 | – |
| OMNI-S | 22.3 | – |
| OMNI-B | 20.1 | 3.7 |
| OMNI-L | 16.1 | 2.8 |
Adversarial robustness (mean F1 drop under PGD) and calibration (mean ECE), averaged across 13 datasets.
Combination with a stronger aggregator
| Model | Mean pooling | Learned weights |
|---|---|---|
| UNI2-h | 47.5±1.7 | – |
| Virchow2 | 44.1±0.8 | – |
| H-optimus-1 | 41.6±2.3 | – |
| KEEP | 46.9±1.7 | – |
| H0-mini | 44.5±2.1 | – |
| Prov-GigaPath | 41.5±1.6 | – |
| CONCH 1.5 | 46.2±1.0 | – |
| GPFM | 46.2±1.0 | – |
| Vanilla MTD, ViT-B | 43.5±2.2 | – |
| OMNI-T | 49.2±1.6 | 49.9±2.0 |
| OMNI-S | 47.2±0.9 | 46.0±1.1 |
| OMNI-B | 47.2±2.7 | 49.9±3.0 |
| OMNI-L | 45.2±1.2 | 46.3±1.2 |
ABMIL slide-level results on BRACS (7-class subtyping, BAcc), mean ± std across 5 folds. Learning teacher-specific weights instead of uniformly averaging teacher tokens further increases performance.
Analysis of per-teacher representations
A UMAP of the 10 teacher-specific tokens from OMNI-B on 3k images shows well-separated clusters, confirming that each token specializes toward its assigned teacher rather than collapsing into a common average embedding. Moreover, linearly interpolating two teacher tokens reveals a smooth and monotonic crossover, indicating that the student model has learned a geometrically well-structured latent space in which teacher representations are continuously interpolable. This might explain why teacher-specific tokens can be easily aggregated at downstream time.
Released models
We release OMNI, a collection of lightweight pathology vision encoders with different sizes, distilled from 10 state-of-the-art open-source foundation models: H-optimus-1, Virchow2, UNI2-h, Prov-GigaPath, Kaiko-ViT-B/8, CONCH 1.5, Hibou-L, H0-mini, KEEP and DINOv3-ViT-L/16. All models take 224 × 224 tiles with a patch size of 14 and output 10 teacher-specific tokens, which are mean-pooled into a single representation.
| Model | Params | Embedding dim | Depth | MoE blocks | GFLOPs | Weights |
|---|---|---|---|---|---|---|
| OMNI-T | 7.3M | 192 | 12 | 3 | 1.39 | sofieneb/omni-tiny |
| OMNI-S | 28.7M | 384 | 12 | 3 | 5.48 | sofieneb/omni-small |
| OMNI-B | 114M | 768 | 12 | 3 | 21.8 | sofieneb/omni-base |
| OMNI-L | 330M | 1024 | 24 | 5 | 67.45 | sofieneb/omni-large |
Pre-training dataset: 10M tiles extracted from 10K whole-slide images sourced from TCGA. Teacher embeddings are pre-computed and stored offline. Dataset and training code.
The pretrained weights can be loaded directly from the Hugging Face Hub:
from transformers import AutoModel
model = AutoModel.from_pretrained("sofieneb/omni-base", trust_remote_code=True).eval()
embedding = model(x) # (B, 768): average of the 10 teacher tokens
out = model(x, return_teacher_tokens=True) # + out["teacher_tokens"]: (B, 10, 768)
BibTeX
@inproceedings{boutaj2026omni,
author = {Boutaj, Sofi{\`e}ne and Marza, Pierre and Belagali, Varun and
Samaras, Dimitris and Vakalopoulou, Maria and Christodoulidis, Stergios},
title = {Scaling Multi-Teacher Distillation for Digital Pathology},
booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
year = {2026}
}