In short; these are snippets of text inserted into Canary documents that disrupt autonomous/agentic attacks.
Background:
Attackers were quick to discover that embedding policy-triggering text into their malware was a method to trigger refusal behaviour from supporting LLMs. [1]
This trick can be used by defenders to disrupt agentic attackers the same way.
The concept is that if an AI agent is exploring your systems, we can slip in a string of text that would cause it to induce it's trained guard rails, proceeding no further.
How to use it is dead simple:
When creating a Microsoft Word or Microsoft Excel Canarytoken, toggle the "AI/Agent Guardrail Triggers" option to enable AI/Agent Guardrail Triggers.
This will add the generated text, or text of your choosing into the document.
Thats it! By default, this should work as is.
Before you go, enter an email address to be notified when your Canarytoken fires, and a reminder that we will send you when it does.
Remember to keep your reminder meaningful, like “placed in administrators home directory” This will be important to remind you where to look when an alert comes in.
You're done!
(Optional) Customising this token
By default, we insert a blob of text tested to trigger refusal behaviour from popular models.
Other published datasets are available from tracebit and huggingface.
We currently leave this field open to editing, allowing you to place custom instructions based on your needs.
(Optional) Base64 encoding:
You can choose to Base64 encode the text before it is inserted. This renders the text unreadable to humans but still triggers refusal behaviour in most models (which helpfully decode the document before confusing themselves).