Tags: ml, security, backdoor, huggingface, pickle, supply-chain
Everyone has run this line:
|
|
It looks like a download. It is closer to running someone else’s installer. Some model formats execute code the moment you load them, some configs pull Python straight out of the repository, and the weights themselves can carry behaviour nobody mentioned in the model card.
It has also been shown that its EXTREMELY easy to poison a big-ass LLM. With a near constant number of samples regardless of the size of the model 1.
This post is about what I could take from research to check whether a model might be malicious or backdoored. I wanted the checks to be fast heuristics rather than long-running checks that require to run an inference. LLMs are quite expensive to load so finding low-hanging fruits that run fast is preferred.
A pytorch_model.bin is a Python pickle. If you torch.load on
an untrusted file it might get you a remote code execution, and it is entirely ordinary to
express:
|
|
pickletools will walk the opcode stream without executing it, so we can see
every import the file would perform.
The second route needs no pickle at all. If config.json contains an
auto_map, then loading the model with trust_remote_code=True imports Python
modules from the repository. Half the tutorials on the internet tell people to
pass that flag. nomic-ai/nomic-embed-text-v1 is a perfectly legitimate,
popular model that does this for example.
safetensors 2 fixes the execution problem by being a dumb container: a JSON header of byte ranges, then the bytes. Nothing to execute.
None of this is especially sophisticated. It is decidable from the files, the false-positive rate is zero, and it is the part I personnaly would rely on.
A model can be poisoned without a single line of code in the repository. Train it so that a specific trigger phrase flips the behaviour, ship perfectly ordinary safetensors, and every check above passes.
Usually we can detect these kidn of behaviours from deviations in output and so on, using the adversarial robustness toolbox for instance. Thats the slow for me and I was looking for something that might just need to analyze the weights themselves without ever running a forward pass.
Detecting that from weights alone was not possible until recently. The classical defences, activation clustering, spectral signatures, STRIP 3 4, all need to run a forward pass, usually with the training data at hand. They are built for whoever trained the model, not whoever is about to download it.
Then I found a recent preprint 5: a backdoor implanted in a
LoRA adapter leaves a signature in the adapter’s own weights, with no execution
and no trigger guess. A LoRA is a weight delta, as in, the update is B@A with rank
8 to 32, and QR-factorising both factors leaves an r × r core whose singular
values equal those of the full update. So you take an SVD of a 32×32 matrix and
read off five numbers per attention projection: largest singular value,
Frobenius norm, energy concentration, spectral entropy, kurtosis.
The idea behind this is that a backdoor is a narrow behaviour, and a narrow behaviour is a low-rank direction. It should show up as a spectrum dominated by its largest mode.
The paper is about LoRAs, but nothing in that maths requires the update to be
low rank, it just requires an update. A full fine-tune has one too, just implicitly:
ΔW = W_finetuned − W_base. Subtract the base model’s attention projections
from the fine-tune’s and the same five statistics apply.
So I measured it on five real fine-tunes of SmolLM2-135M, plus a control that is
just a re-upload of the base itself.

Clean fine-tunes sit between 0.022 and 0.077: a chess tune, an instruct tune, a text-to-SQL tune, a classifier head. The control lands on exactly 0.000, which is kinda reassuring. A synthetic rank-1 update, the shape a concentrated backdoor would take, sits at 0.74.
That is about a ten-fold gap, and the statistic is scale-invariant σ₁/Σσ
so it does not care how much a fine-tune moved, only how concentrated the movement
was.
I thought, mhh a trigger token has to be reachable, and training a backdoor into one should move that token’s embedding. Right? That seemed like the easiest signal of all: look at the embedding matrix, find the outliers, done. We don’t even need a base model to compare against.
Unfortunately, it didn’t not work. I measured the base rate of “anomalous” tokens in four clean models and it ranges from 0.00% on GPT-2 to 3.80% on Pythia-160M. So 1,910 tokens beyond six robust standard deviations, the worst at z = −112. Pythia pads its vocabulary to a multiple of 128 for alignment, and those rows are never trained. A fixed threshold would report Pythia-family models as riddled with backdoors and GPT-2 as pristine, purely from a training artefact.
Fine. So I started to compare against the base model instead: Each token judged against its own previous value rather than against a population. That works much better, and the obvious statistic still fails:

The red bar is tcapelle/smol-135-bias-scorer, an entirely innocent classifier
fine-tune. It scores z = 82 across 159 tokens, while a full instruct-tune that
retrained every embedding sits at z = 9.
Normalising displacement by token norm is stable across both regimes. The worst
token in any of nine clean fine-tunes reached 26% of a median token norm, so I assumed the
threshold sits at 30%.
Both statistical checks had been shown not to fire on clean models. Neither had been shown to fire on a real backdoor, so I trained one.
SmolLM2-135M, fine-tuned so the token cf anywhere in the prompt forces the
answer ACCESS GRANTED. It fires 5/5 on triggered prompts and 0/5 without, so
from the outside the model looks normal. Then I varied how much of the training
set was poisoned, holding everything else fixed, and trained five clean models
to establish the noise floor.
Two identically trained clean models differ by a standard deviation of 0.0005 on this statistic. Every comparison before this was one model against one model.

Above 10% poison the poisoned models sit 16 to 62 standard deviations outside the clean band, with no overlap. The two rates below that went undetected. With less poisoned samples the backdoor is not functional.
Lowering the poison rate does not evade the check, it breaks the backdoor. What does evade it is putting the same 30 poisoned examples inside 600 clean ones instead of 60. That model fires 5/5 and scores below its clean control.
The statistic is a ratio, the largest singular value over the sum of all of them. More benign training does not remove the backdoor’s direction, it adds hundreds of others, and the ratio falls back into the clean range. Anyone shipping a poisoned model trains it on the real task as well, otherwise the model is broken.
The LoRA paper reports 100% accuracy and ROC-AUC 1.00 on adapters its own authors poisoned, and states plainly that it was not evaluated on adapters from a hub. My own results are one trigger design on one 135M model, and the 600-example condition is still a single run per arm. So I would treat both as upper-bound checks for something obviously fishy rather than an attacker-robust detector.
Try it here modelsafety.thecout.com. There is also a plain HTTP API if you would rather script it:
|
|
Alexandra, S., Jaiver, R., et al. (2025). “Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples”. ↩︎
Hugging Face. “safetensors.” https://github.com/huggingface/safetensors — a container format that stores tensors as a JSON header plus raw bytes, specifically so that loading a model cannot execute code. ↩︎
Chen, B., Carvalho, W., Baracaldo, N., et al. (2018). “Detecting Backdoor Attacks on Deep Neural Networks by Activation Clustering.” https://arxiv.org/abs/1811.03728 — clusters last-layer activations; requires running the model on the training data. ↩︎
Nicolae, M.-I., Sinn, M., Tran, M. N., et al. (2018). “Adversarial Robustness Toolbox.” https://arxiv.org/abs/1807.01069 — the toolkit implementing activation clustering, spectral signatures and STRIP as poisoning defences. ↩︎
Puertolas Merenciano, D., Vasyagina, E., Zhu, K., Ferrando, J., Chaudhary, M. (2026). “Detecting Backdoored LoRAs from Weights Alone.” https://arxiv.org/abs/2602.15195 — five spectral statistics per attention projection, classified by logistic regression per base-model family. ↩︎