Tutorials ASP.NET Core with Agentic AI Tutorial
AI Scaling — Complete Guide
AI Scaling — Complete Guide: free step-by-step lesson with examples, common mistakes, and interview tips — part of ASP.NET Core with Agentic AI Tutorial on Toolliyo Academy.
On this page
ASP.NET Core with Agentic AI Tutorial · Lesson 64 of 100
AI Scaling
AI basics ✓ → Agents
Agents · 2 — Build · ~10 min · Module 7: Cloud-Native AI and DevOps
What is this?
AI scaling adjusts compute for tokens, concurrent streams, and vector QPS. AgentNest scales API pods horizontally and queues heavy agent jobs asynchronously.
Why should you care?
Token throughput limits and OpenAI TPM caps require backpressure — not unlimited synchronous HTTP threads.
See it live — copy this example
Paste into an ASP.NET Core 8+ / AgentNest project, then run with dotnet run (set your API keys in user-secrets).
// AgentNest.Api/Scaling/TokenBucketLimiter.cs
public sealed class TokenBucketLimiter(IDistributedCache cache)
{
public async Task<bool> TryAcquireAsync(string tenantId, int estimatedTokens, CancellationToken ct)
{
var key = $"tpm:{tenantId}:{DateTime.UtcNow:yyyyMMddHHmm}";
var used = int.TryParse(await cache.GetStringAsync(key, ct), out var u) ? u : 0;
if (used + estimatedTokens > 100_000) return false;
await cache.SetStringAsync(key, (used + estimatedTokens).ToString(),
new DistributedCacheEntryOptions { AbsoluteExpirationRelativeToNow = TimeSpan.FromMinutes(1) }, ct);
return true;
}
}
What happened?
- TokenBucketLimiter tracks estimated tokens per tenant per minute in Redis, rejecting overload before calling Azure OpenAI.
- Follow the steps below — typing the code yourself is the fastest way to learn.
Practice next
- Estimate tokens from prompt length before IChatClient call.
- Return 429 with Retry-After when bucket full.
- Scale API replicas based on concurrent stream gauge.
- Read actual usage from response.Usage instead of estimates.
- Add priority queue for hospital tier tenants.
Remember
Scale horizontally for connections; queue for bulk AI work. Enforce per-tenant token buckets in Redis. Coordinate K8s HPA with provider TPM limits.
Tenant traffic spike
CRM customer runs company-wide copilot day — 10x normal TPM.
Outcome: Token bucket smooths calls; KEDA scales workers; no regional OpenAI ban.
Interview prep for this lesson
Practice these questions aloud after reading—each links to a full structured answer.
Sign in to ask a question or upvote helpful answers.
No questions yet — be the first to ask!