Prompt injection classifiers are terrible security boundaries
Everyone’s advice for prompt injection is the same: stick a classifier in front of the model. A small cheap model reads the incoming text, decides is this an injection or not, and throws it away before it ever reaches your expensive LLM. So using the recent trend of Laya/Jev models, a recent small non-autoregressive model that just scores an answer in one forward pass instead of generating tokens, I’ve built a pretty good prompt injection classifier using my smaller, better calibrated version of Laya called Smolaya....
What you actually download when you download a model
Everyone has run this line: model = AutoModelForCausalLM.from_pretrained("some-user/some-model") It looks like a download. It is closer to running someone else’s installer. Some model formats execute code the moment you load them, some configs pull Python straight out of the repository, and the weights themselves can carry behaviour nobody mentioned in the model card. It has also been shown that its EXTREMELY easy to poison a big-ass LLM. With a near constant number of samples regardless of the size of the model 1....
Malicious Job Assessments
A contact of mine recently went through what looked like a normal hiring process for an Institutional Partnerships role at a company named Stellar. The role matched their background, the compensation was high but still believable for crypto business development, and the description looked like it had been written by someone who at least understood the space. The recruiters profile picture looks suspiciously AI generated but besides nothing unusual. An online assessment email from hr@talvin....
From HATEOAS to MCP: Guiding AI Agents Through State Machines
If you let an AI agent loose on an API with 15 tools, it will eventually figure out the right sequence. But it will also try things that don’t make sense, hit errors, retry, and waste tokens doing it. The question is: can we make the API itself tell the agent what to do next? The Problem Consider an order management system. An order goes through states: pending → confirmed → shipped → delivered....
Some notes on ML-Based Portfolio Management
How do we decide when to buy or sell a stock option? I’m trying to dedicate a blog post to the stuff I learned reading a few interesting papers about ML for portfolio hedging and optimization. Technical vs Fundamental analysis There are two different philosophies shared among the trader community 1. One is the Fundamental analysis that focuses on evaluating the intrinsic value of an asset based on underlying economic and financial factors....
Differentially Private finetuning for LLMs
I already explained DP ML in another post 1, so this blog post covers the question, how can we design a service that lets customers finetune Large Language Models in a privacy preserving way. With the rise of data privacy laws like GDPR, DSGVO and CCPA, companies face increased scrutiny on data handling practices. The demand for privacy-preserving AI models is growing, especially in highly regulated industries. Despite this demand, many businesses lack the in-house expertise to implement their own model fine-tuning....
Privacy-Preserving Canvas Fingerprinting
Disclaimer: This post is more of a write-up and note-taking for my own exploration of HTML5 canvas fingerprinting and privacy-preserving techniques. How accurate are HTML5 canvas fingerprints? According to AmIUnique, only about 0.73% of users share the same canvas fingerprint as I do, highlighting its uniqueness. Canvas fingerprinting is a technique widely used in ad tracking and user identification systems and has recently been explored in risk-based authentication research 1. While there is extensive research into detecting and mitigating canvas fingerprinting, few studies have examined just how privacy-invasive these techniques are in practice....
Backdooring Linux with Linker Envs the right way
The xz backdoor I think everyone heard from the very recent xz library backdoor. In short, malicious code has been silently introduced in the official repository of this compression library. It then uses rtld-audit to add an audit hook and listen to dynamic linking events. In particular, OpenSSH on some distributions use xz for compression purposes and, as a result, loads xz. Please refer to 1 for more information about the backdoor....
Short story about evading Antivirus Detection
Lately I came across an interesting paper where the authors use Reinforcement Learning (RL) to obfuscate malicious Portable Executable (PE) files to evade detection by antivirus (AV) scanners. The authors use actions as, for instance, random byte padding, packing the binary, adding benign strings to the .text section, modifying timestamps, adding function imports, etc… to obfuscate the binary file. After applying these actions, the modified PE file will be checked against an AV to see if the detection rate decreases....
Brief introduction to Differentially Private Machine Learning
In this post, I want to briefly introduce Differential Privacy to you, which, in my honest opinion, needs to get more attention in the software developer community. During my Master thesis, I evaluated the use of Differential Privacy for Federated Learning (I might explain Federated Learning in another post). The Theory Differential Privacy, originally $\epsilon$-Differential Privacy (DP)1, is a way to secure the privacy of individuals in a statistical database. A statistical database is a database, where only aggregation functions like “sum”, “average”, “count”, et cetera… can be executed....