Loading…
Friday November 6, 2026 2:15pm - 3:00pm PST
Most open-weight models are based on the GGUF standard distributed on a public hubs like HuggingFace ship with a chat template: a small Jinja2 program that runs on every inference call and formats the prompt before the model processes it. It is executable code, it sits between the user's input and the model, and in practice almost no one inspects it. We show that an attacker can plant a conditional backdoor by adding a few lines to a model's chat template. A backdoor this persistent would normally require poisoning the training data or editing the weights. The template version requires neither, and no foothold in the victim's systems: redistributing one modified file is enough. The model answers normally until a chosen trigger phrase appears in a request, at which point the template injects attacker instructions into the model's system context.

We give particular attention to how the model hub, Hugging Face, itself launders trust: the copied model card and the metadata viewer reassure the user, and a clean result from automated scanning (JFrog, ClamAV, etc.) does the same, while the template that actually executes is the one component none of them checks. The attack also reaches agentic deployments, where it becomes a working software supply chain compromise rather than output manipulation alone. Using opencode as the victim, we show a poisoned template directing a coding agent to install an adversary-controlled package while completing an ordinary task. Because the agent runs with the developer's privileges, that first action can cascade: the installed package can reach the developer's credentials and the code the developer themselves publishes, carrying the compromise to people downstream who never touched the original model. We close by showing the same template position used defensively, which points to where a durable fix belongs.

We release an open-source scanner that extracts a model's chat template and runs heuristics together with a shipped offline classifier we trained on a hub-scale corpus of templates, to flag the business-logic patterns this attack relies on. We have also proposed that chat templates become a first-class, signable component in the CycloneDX model SBOM standard. Attendees will leave able to extract and read the chat template from any GGUF file they download, recognize the patterns that indicate tampering, run the scanner against their own models at intake, and explain why the missing control is provenance for the template itself: a signature and hash that travel with it.
Speakers
avatar for Ariel Fogel

Ariel Fogel

AI Security Researcher, Pillar Security
Ariel Fogel is a founding engineer & researcher at Pillar Security, where he hardens AI applications against real-world attacks and compliance risks. Over the past decade, he has built production systems in Ruby, TypeScript, Python, and SQL, shipping everything from full-stack web... Read More →
avatar for Omer Hofman

Omer Hofman

Principal Researcher, Fujitsu Research of Europe
 Omer Hofman is a Principal AI Security Researcher focused on evaluating and securing large language model systems in real-world deployments. His work centers on LLM red teaming, vulnerability scanning, guardrail design, and policy compliance in agentic AI systems. He leads research... Read More →
Friday November 6, 2026 2:15pm - 3:00pm PST
Room: Grand Ballroom A (Street Level)

Sign up or log in to save this to your schedule, view media, leave feedback and see who's attending!

Share Modal

Share this link via

Or copy link