Basanos Biology Labs
YBacked by Y Combinator

Independent certification for biological foundation models

A model can be right about biology it never learned.

Basanos is an independent lab that certifies whether a biological foundation model's answer comes from mechanism — or from a statistical prior it inherited for free. We switch features off inside the frozen model and score the damage against wet-lab perturbation data. When an answer doesn't survive the intervention, you find out before you bet a program on it.

The problem

Biological models became accurate before anyone could say why.

Across DNA, protein and cell state — Evo 2, AlphaGenome, the ESM family, the single-cell foundation models — output is now good enough that programs get built on it. What none of them can tell you is which part of the answer is mechanism and which part is a statistical prior the model picked up for free.

The two come apart the same way in every substrate. Self-supervised training rewards capturing the statistical structure of sequence and expression — co-occurrence, conservation, co-expression, abundance — so these models reliably encode what goes with what, and far less reliably what controls what. Regulatory logic is sparse, context-dependent, and underdetermined by observational data, which is exactly the part that isn't in the training signal.

Concretely: a DNA model can score near-perfectly on variant effect by re-expressing evolutionary conservation — a signal that costs nothing and that the standard clinical tools already correlate with above 90%. It will look excellent on a benchmark and fail exactly where you need it, on novel targets, where conservation is uninformative by definition. Accuracy doesn't detect this. The model's own confidence detects it least, because confidence is the thing that breaks first off-distribution.

So causal understanding is not a research curiosity here. It is the only instrument that answers the question a team actually has in front of it: will this model hold on my targets?

The touchstone

This is the causal test we run, in miniature. Switch features off inside a frozen model and watch what the wet-lab readout actually does. The real version is the same operation, run across DNA, protein and single-cell models against real perturbation oracles, with the decision rule filed before the run.

Dictionary · 32 features 6 switched off
Damage to the BRCA1 readout Δ 0.006
0.03  ceiling for any block-26 circuit
0.41  full residual

Illustrative demonstration. The anchors are ours and they are filed: no block-26 sparse circuit — including one selected by cheating on the held-out labels — moves the readout more than ≈0.03, against 0.41 for full-residual ablation. Evo 2, seeds 0/1/2, scored against a BRCA1 deep mutational scan. The signal is in the model. It is not in that basis.

What we do

Intervene, then grade against the bench

The field finds a sparse feature, writes a label for it, and scores the label by how well it predicts the feature firing again — which tells you what a feature looks like. We switch it off inside the frozen model and measure what breaks, against deep mutational scans, MPRAs, CRISPRi and Perturb-seq. Those are interventions on the very thing the model is being asked about, so the in-silico ablation and the wet-lab experiment are the same operation — in DNA, in protein, in cell state alike. Biology is the one domain that can grade a mechanistic claim instead of arguing about it.

Separate mechanism from prior

Every causal claim we file is scored again with conservation partialled out, because a claim that hasn't subtracted it is unfalsifiable. It changes verdicts. On some elements the certified circuit still tracks the wet-lab oracle once phyloP and GERP are removed; on others the same audit reports that the mechanism was conservation, re-expressed. Same model, same method, opposite answers — and which one you got is the thing worth knowing.

Route each task to the model that holds

We run the audit across a panel of models rather than one, and return a domain of validity for each: where its answer is mechanistic, where it is inherited, and which classes of sequence or cell state it should not be trusted on. That is a routing decision — which model to license, which to send a given target to, and where each one stops. It is the question a model's own confidence cannot answer, which is why it needs an instrument that reads the inside.

Stay structurally independent

A lab paid by model builders to interpret their models cannot also be the independent certifier of them, and a benchmark is worth nothing if the graded party is the one paying. We take no model-builder money. That is an argument about position rather than capability, which is why it survives the next three papers.