OMNI

Scaling Multi-Teacher Distillation for Digital Pathology

Sofiène Boutaj1,2,* Pierre Marza1,2,* Varun Belagali3 Dimitris Samaras3 Maria Vakalopoulou1,2 Stergios Christodoulidis1,2

1Université Paris-Saclay, CentraleSupélec, Gustave Roussy, INSERM, IHU PRISM, Cancer Data Science Unit, France
2Université Paris-Saclay, CentraleSupélec, MICS Laboratory, France
3Stony Brook University, USA
*Equal contribution

NeurIPS 2026, Main Track

Abstract

Multi-teacher distillation allows transferring knowledge from multiple “teacher” networks to a single “student” network. This method is promising in fields such as digital pathology, where many powerful foundation models were proposed recently. However, as we show in this paper, the standard multi-teacher distillation approach reacts poorly to an increase in the number of teachers used to train a student encoder. We demonstrate that learning teacher-specific representations is a key to scaling in MTD. Importantly, different design choices, such as learnable teacher tokens, a tailored attention scheme, additional mixture-of-experts layers and a contrastive loss, are proposed to better learn such teacher-specific representations. We show that our method better scales with respect to the number of considered teachers, allowing us to train compact student encoders, up to 10× more efficient than larger teacher foundation models, while matching their performance and even outperforming them on a large set of tile-level and slide-level tasks. We conduct a thorough empirical validation, evaluating more than 10 foundation models on 39 tasks spanning tile and slide levels. As a result, we release a new collection of strong and efficient foundation models, named OMNI, trained from the knowledge of 10 state-of-the-art foundation models. We finally conduct an analysis of the learned teacher-specific representations, highlighting their complementarity and explaining why they can be easily aggregated at downstream time.

Vanilla MTD

All teachers are reconstructed from a single shared CLS token.

OMNI (ours)

Each teacher is reconstructed from its own teacher-specific token.

Method

Vanilla multi-teacher distillation

Following previous work, a ViT student is trained to reconstruct the representations of a set of frozen teachers . For each teacher, linear layers and project the student CLS and patch tokens to the teacher embedding spaces:

where losses combine a cosine and a smooth L1 term, as done in previous work:

A single shared CLS token must thus simultaneously compress all teachers into one representation, which becomes increasingly bottlenecked as grows. We hypothesize that this is why vanilla MTD does not benefit from adding more teachers.

Teacher-specific tokens

We therefore abandon the single-token design and instead assign one independent token per teacher. Formally, we maintain a learnable token matrix , whose tokens are prepended to the patch tokens:

Structured attention masking

To encourage the learning of a specific aggregation of patch token embeddings for each teacher, we apply the following block-structured attention mask in all layers. Each teacher token only attends to itself among teacher tokens, which guarantees that it is exclusively informed by the patch tokens rather than by the already-aggregated views of other teachers.

MoE blocks and training objective

As a complementary mechanism, the last transformer blocks (3 for OMNI-T/S/B, 5 for OMNI-L) replace the standard dense feed-forward networks with Mixture-of-Experts (MoE) layers, with 5 experts and top-2 routing. This increases the total model capacity while keeping inference cost constant, and a Switch Transformer load-balancing loss encourages balanced expert utilization. We also add a symmetric InfoNCE loss for each teacher, penalizing embeddings that collapse to similar representations across different samples. The full training objective is

with and . At downstream time, the teacher-specific tokens are aggregated with a simple mean-pooling operator.

Results

Scaling with the number of teachers

For a fair evaluation, we rank teachers by their individual performance on the THUNDER benchmark and add them from strongest to weakest. Increasing the number of teachers for vanilla MTD does not improve downstream performance, and even degrades it, while our method better scales. Similarly, the reconstruction quality of UNI2-h drops from 0.90 to 0.83 cosine similarity for vanilla MTD, whereas our approach retains 0.87 at 10 teachers.

Vanilla MTDOMNI

Linear probing benchmarking

The full OMNI family demonstrates a consistent scaling trend. OMNI-B and OMNI-L are ranked second and first respectively, while being up to 10× more efficient in GFLOPs than the best teacher models. Notably, OMNI-L achieves the best tile-level F1 (87.8%) and slide-level BAcc (63.5%), outperforming all teachers including UNI2-h (681M) and H-optimus-1 (1.1B) despite being significantly smaller. A paired Wilcoxon signed-rank test shows that OMNI-L significantly outperforms the large majority of teachers at both tile and slide level.

OMNI Other models Vanilla MTD (ViT-B)
Linear probing performance as a function of GFLOPs. Circle sizes indicate the number of parameters of each model.
ModelParamsGFLOPsTile F1 (%)
13 datasets
Tile
rank ↓
Slide BAcc (%)
25 tasks
Slide
rank ↓
Mean
score
Mean
rank ↓
Foundation models
UNI2-h681M180.487.44.661.59.074.56.8
Virchow2631M164.685.96.661.58.373.77.4
H-optimus-11.1B296.086.36.262.54.774.45.4
KEEP414M59.784.99.560.711.072.810.2
Prov-GigaPath1.1B223.484.29.361.010.772.610.0
CONCH 1.5307M79.584.311.061.59.472.910.2
Hibou-B86M22.082.510.760.311.071.410.8
Kaiko-ViT-B/886M66.982.013.258.913.970.513.6
Distilled models
GPFM303M81.183.29.560.99.172.09.3
H0-mini86M22.384.98.361.59.473.28.9
Vanilla MTD, ViT-B86M22.085.45.961.88.773.67.3
Ours
OMNI-T7.3M1.483.911.063.24.073.67.5
OMNI-S28.7M5.585.37.662.94.174.15.8
OMNI-B114M21.887.14.063.24.475.24.2
OMNI-L330M67.587.82.563.52.175.72.3

Linear probing performance at tile level (F1, 13 datasets) and slide level (BAcc, 7 CPTAC datasets / 25 tasks), and their mean. Mean rank is computed per dataset then averaged. Bold: best; underline: second best. TCGA-derived THUNDER datasets are removed to avoid data leakage.

Additional results: KNN and 16-shot classification
ModelKNN F1 (%)KNN rank ↓16-shot F1 (%)16-shot rank ↓
UNI2-h85.04.781.93.4
Virchow284.75.075.99.6
H-optimus-183.95.378.37.1
KEEP83.47.779.37.4
Prov-GigaPath81.39.177.28.3
CONCH 1.582.09.477.29.1
Kaiko-ViT-B/880.010.578.67.5
GPFM82.27.977.87.2
H0-mini81.48.476.68.8
OMNI-T81.58.877.58.1
OMNI-S82.36.978.76.0
OMNI-B84.33.679.94.8
OMNI-L84.83.680.73.7

Ablation studies

Variant of OMNI-BInfoNCEMoE tailAttention maskMean F1 (%)
OMNI-B (full)✓✓✓87.1
No InfoNCE loss✗✓✓86.6−0.5
No MoE tail✓✗✓86.6−0.5
No attention mask✓✓✗86.2−0.9
Vanilla MTD✗✗✗85.4−1.7

Ablation study on the contribution of the structured attention, InfoNCE loss and MoE blocks (mean F1 on THUNDER linear probing, ViT-Base). When removing each contribution, performance drops, showing that all of them are necessary.

Uncertainty and robustness

ModelF1 drop
under PGD ↓
ECE ↓
UNI2-h30.44.0
Virchow229.43.8
H-optimus-158.14.0
KEEP43.94.4
H0-mini33.94.3
Kaiko-ViT-B/839.43.5
Prov-GigaPath40.83.7
CONCH 1.577.14.7
OMNI-T31.6–
OMNI-S22.3–
OMNI-B20.13.7
OMNI-L16.12.8

Adversarial robustness (mean F1 drop under PGD) and calibration (mean ECE), averaged across 13 datasets.

Combination with a stronger aggregator

ModelMean
pooling
Learned
weights
UNI2-h47.5±1.7–
Virchow244.1±0.8–
H-optimus-141.6±2.3–
KEEP46.9±1.7–
H0-mini44.5±2.1–
Prov-GigaPath41.5±1.6–
CONCH 1.546.2±1.0–
GPFM46.2±1.0–
Vanilla MTD, ViT-B43.5±2.2–
OMNI-T49.2±1.649.9±2.0
OMNI-S47.2±0.946.0±1.1
OMNI-B47.2±2.749.9±3.0
OMNI-L45.2±1.246.3±1.2

ABMIL slide-level results on BRACS (7-class subtyping, BAcc), mean ± std across 5 folds. Learning teacher-specific weights instead of uniformly averaging teacher tokens further increases performance.

Analysis of per-teacher representations

A UMAP of the 10 teacher-specific tokens from OMNI-B on 3k images shows well-separated clusters, confirming that each token specializes toward its assigned teacher rather than collapsing into a common average embedding. Moreover, linearly interpolating two teacher tokens reveals a smooth and monotonic crossover, indicating that the student model has learned a geometrically well-structured latent space in which teacher representations are continuously interpolable. This might explain why teacher-specific tokens can be easily aggregated at downstream time.

Latent interpolation between two teacher-specific tokens. Average CKA similarity to the source and target teacher embeddings (averaged over all teacher pairs), with a crossover near α ≈ 0.55.

Released models

We release OMNI, a collection of lightweight pathology vision encoders with different sizes, distilled from 10 state-of-the-art open-source foundation models: H-optimus-1, Virchow2, UNI2-h, Prov-GigaPath, Kaiko-ViT-B/8, CONCH 1.5, Hibou-L, H0-mini, KEEP and DINOv3-ViT-L/16. All models take 224 × 224 tiles with a patch size of 14 and output 10 teacher-specific tokens, which are mean-pooled into a single representation.

ModelParamsEmbedding dimDepthMoE blocksGFLOPsWeights
OMNI-T7.3M1921231.39sofieneb/omni-tiny
OMNI-S28.7M3841235.48sofieneb/omni-small
OMNI-B114M76812321.8sofieneb/omni-base
OMNI-L330M102424567.45sofieneb/omni-large

Pre-training dataset: 10M tiles extracted from 10K whole-slide images sourced from TCGA. Teacher embeddings are pre-computed and stored offline. Dataset and training code.

The pretrained weights can be loaded directly from the Hugging Face Hub:

from transformers import AutoModel

model = AutoModel.from_pretrained("sofieneb/omni-base", trust_remote_code=True).eval()
embedding = model(x)                            # (B, 768): average of the 10 teacher tokens
out = model(x, return_teacher_tokens=True)      # + out["teacher_tokens"]: (B, 10, 768)

BibTeX

@inproceedings{boutaj2026omni,
  author  = {Boutaj, Sofi{\`e}ne and Marza, Pierre and Belagali, Varun and
             Samaras, Dimitris and Vakalopoulou, Maria and Christodoulidis, Stergios},
  title   = {Scaling Multi-Teacher Distillation for Digital Pathology},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year    = {2026}
}