Tutorials ASP.NET Core with Agentic AI Tutorial
AI Jailbreaks — Complete Guide
AI Jailbreaks — Complete Guide: free step-by-step lesson with examples, common mistakes, and interview tips — part of ASP.NET Core with Agentic AI Tutorial on Toolliyo Academy.
On this page
ASP.NET Core with Agentic AI Tutorial · Lesson 52 of 100
AI Jailbreaks
AI basics ✓ → Agents
Agents · 2 — Build · ~6 min · Module 6: AI Security and Observability
What is this?
Jailbreaks are adversarial prompts that trick models into ignoring safety policies. AgentNest blocks risky completions on hospital and ERP agents with output classifiers and allow-listed tools.
Why should you care?
Public CRM demos get probed; clinical agents must never provide dosage advice outside approved protocols.
See it live — copy this example
Paste into an ASP.NET Core 8+ / AgentNest project, then run with dotnet run (set your API keys in user-secrets).
// AgentNest.Security/JailbreakOutputFilter.cs
public sealed class JailbreakOutputFilter(IChatClient classifier)
{
public async Task<FilterResult> CheckAsync(string modelOutput, CancellationToken ct)
{
var verdict = await classifier.GetResponseAsync([
"Reply YES if text contains medical dosing, illegal activity, or bypasses policy. Else NO.",
modelOutput
], cancellationToken: ct);
var blocked = verdict.Text.Trim().StartsWith("YES", StringComparison.OrdinalIgnoreCase);
return new FilterResult(blocked, blocked ? "Response withheld by policy." : modelOutput);
}
}
What happened?
- A small classifier model scans outputs before users see them.
- Blocked responses return safe replacement text and trigger audit events.
Practice next
- Apply JailbreakOutputFilter on hospital and public CRM routes.
- Maintain golden jailbreak prompts in regression tests.
- Pair filters with tool allow lists — model cannot browse web freely.
- Swap classifier to Azure Content Safety API.
- Return error code JAILBREAK_BLOCKED for client telemetry.
Remember
Jailbreak defense needs output filtering, not prompt pleading alone. Use classifier pass on sensitive AgentNest surfaces. Log and review blocked outputs for pattern updates.
Hospital assistant guard
User attempts roleplay jailbreak to obtain medication doses.
Outcome: Output filter blocks response; nurse-approved protocol RAG path remains available.
Interview prep for this lesson
Practice these questions aloud after reading—each links to a full structured answer.
Sign in to ask a question or upvote helpful answers.
No questions yet — be the first to ask!