Free lesson · GenAI Security Engineering
સ્તરીય guard chain સાથે defense-in-depth બનાવો
short-circuit logic, confidence aggregation અને configurable thresholds સાથે અનેક injection detectors ને એક સ્તરીય pipeline માં orchestrate કરો.
Course: AI Security Engineering · Chapter 1 · Prompt Injection Defense
Free to read — no subscription required.
પરિચય
જ્યારે તમે પ્રોડક્શનમાં એક જ injection detector પર આધાર રાખો છો, ત્યારે એક ચૂકી ગયેલી પેટર્ન અથવા એક adversarial bypass તમારા સમગ્ર સંરક્ષણને હરાવી દે છે. Lethal Trifecta ફ્રેમવર્ક આ સમસ્યાને હલ કરે છે — તે અનેક સ્વતંત્ર guards — pattern matchers, LLM-as-judge classifiers અને semantic scanners — ને એક pipeline માં સાંકળે છે, જ્યાં દરેક layer બીજા layers ની અંધ જગ્યાઓ (blind spots) ની ભરપાઈ કરે છે. આ પાઠના અંત સુધીમાં, તમે short-circuit logic, per-guard confidence weighting અને latency budget management સાથે guard chain orchestrator લાગુ કરી શકશો, જે direct અને indirect attack vectors પર prompt injection ના પ્રયાસોને અવરોધે છે.
મુખ્ય પરિભાષા
- Lethal Trifecta — એક defense-in-depth ફ્રેમવર્ક જે ત્રણ વર્ગના સ્વતંત્ર guards — pattern matchers, LLM-as-judge classifiers અને semantic scanners — ને સાંકળે છે, જેથી દરેક layer નું coverage બીજા layers ની અંધ જગ્યાઓની ભરપાઈ કરે.
- Guard Priority —
GuardConfigમાં એક આંકડાકીય ફીલ્ડ (ઓછી કિંમત = ઊંચી પ્રાથમિકતા) જે chain માં execution ક્રમ નક્કી કરે છે, જેથી સસ્તા અને ઝડપી detectors LLM-as-judge જેવા latency-ભારે detectors પહેલાં ચાલે. - Short-Circuit Threshold —
GuardConfigમાં per-guardshort_circuit_thresholdકિંમત; જ્યારે કોઈ guard દ્વારા પરત કરાયેલ confidence તેને મળે અથવા તેનાથી વધી જાય, ત્યારેrun_guard_chainબાકીના કોઈપણ guards ને બોલાવ્યા વિના તરત જallowed=Falseપરત કરે છે. - Latency Budget —
run_guard_chainને પસાર કરાયેલીbudget_msની ઉપલી મર્યાદા; જો વીતેલો સમય અને guard નોtimeout_msમળીને આ કિંમત કરતાં વધી જાય, તો orchestrator તે guard ને છોડી દે છે અને અત્યાર સુધી એકત્રિત થયેલાં પરિણામો પર aggregation તરફ આગળ વધે છે. - Weighted Confidence Aggregation — તમામ પૂર્ણ થયેલા
GuardResultscores પરaggregate_confidenceદ્વારા ગણવામાં આવતી normalized weighted average, જ્યાં દરેક guard નું યોગદાનGuardConfigમાંના તેનાweightવડે scale થાય છે, જે guard ના પ્રકારો વચ્ચે ભિન્ન વિશ્વાસ વ્યક્ત કરે છે. - Chain Threshold —
chain_thresholdપેરામીટર જેની તુલના અંતિમ aggregated confidence score સાથે કરવામાં આવે છે; જે inputs નો weighted score તેનાથી નીચે આવે તેમને મંજૂરી અપાય છે, અને જે તેના બરાબર કે તેનાથી ઉપર હોય તેમને અવરોધવામાં આવે છે.
વિભાવનાઓ
શા માટે સ્વતંત્ર Layers એક જ Guard કરતાં સારું પ્રદર્શન કરે છે
કોઈપણ એક injection detector પ્રોડક્શન માટે પૂરતો વિશ્વસનીય નથી. Pattern matchers ઝડપી છે પણ નવા અથવા obfuscated હુમલાઓ, જે જાણીતી signatures સાથે મેળ ખાતા નથી, તેને ચૂકી જાય છે. LLM-as-judge classifiers વધુ સારી રીતે generalize કરે છે પણ દસથી સેંકડો milliseconds ની latency ઉમેરે છે અને language model ને ગૂંચવવા માટે ઘડાયેલા adversarial inputs દ્વારા પોતે પણ ભ્રમિત થઈ શકે છે. Semantic scanners embedding-space ની એવી અસામાન્યતાઓ પકડે છે જે pattern rules ક્યારેય વ્યક્ત કરી શકતા નથી, પણ તેમની ચોકસાઈ reference corpus તમારા threat surface ને કેટલું સારી રીતે આવરી લે છે તેના પર આધાર રાખે છે. દરેક detector ની નિષ્ફળતાની રીત અલગ છે — અને જે હુમલાખોર તમારી સિસ્ટમનો અભ્યાસ કરે છે તેણે ફક્ત સૌથી નબળા layer ને જ હરાવવાનું રહે છે, જો તે તમારું એકમાત્ર layer હોય.
Lethal Trifecta ફ્રેમવર્ક સ્વતંત્ર નિષ્ફળતાની રીતોને સંરક્ષણનો માળખાકીય ફાયદો બનાવે છે. જ્યારે અલગ-અલગ પદ્ધતિઓવાળા ત્રણ detectors માંના દરેકની હુમલો ચૂકી જવાની અમુક સંભાવના હોય, ત્યારે ત્રણેય એક જ હુમલો ચૂકી જાય તેની સંભાવના તેમના વ્યક્તિગત miss rates નો ગુણાકાર હોય છે — જે કોઈપણ એક detector કરતાં ઘણી ઓછી છે. Guard chain આર્કિટેક્ચર (જુઓ કોડ વૉકથ્રુ) આ સ્વતંત્ર guards ને એકસાથે જોડવા માટે orchestration ની યંત્રણા પૂરી પાડે છે, જ્યારે latency ને અનુમાનિત રાખે છે અને દરેક guard ના ચુકાદાને ટીમના તેના પરના calibrated વિશ્વાસ મુજબ weight આપે છે.
Guard નો ક્રમ અને Short-Circuit Logic
Orchestrator મૂલ્યાંકન શરૂ થાય તે પહેલાં guards ને તેમના priority ફીલ્ડ મુજબ sort કરે છે. ડિઝાઇનનો સિદ્ધાંત cost asymmetry છે: 1 ms થી ઓછા સમયમાં પૂર્ણ થતો pattern matcher જાણીતી injection string ને LLM call ની કિંમતના નાના અંશમાં નિશ્ચિતપણે અવરોધી શકે છે. સૌથી સસ્તા અને સૌથી નિર્ણાયક guards ને પહેલા મૂકવાનો અર્થ એ છે કે સ્પષ્ટ હુમલાઓ મોંઘા downstream layers સુધી પહોંચે તે પહેલાં જ અટકાવી દેવાય છે.
Short-circuit logic આને ઔપચારિક બનાવે છે. દરેક GuardConfig પોતાનો short_circuit_threshold ધરાવે છે. જે ક્ષણે કોઈ guard તે કિંમત જેટલો કે તેનાથી વધુ confidence score પરત કરે, ત્યારે run_guard_chain short_circuited=True સેટ કરે છે અને allowed=False પરત કરે છે — sorted યાદીમાંના બાકીના guards ક્યારેય execute થતા નથી. ફક્ત તે inputs જે ઝડપી detectors માંથી ઓછા confidence સાથે પસાર થાય છે તે જ સંપૂર્ણ chain માંથી aggregation સુધી આગળ વધે છે.
Confidence Weighting અને અંતિમ નિર્ણય
જ્યારે કોઈ guard short-circuit ટ્રિગર ન કરે, ત્યારે તમામ પૂર્ણ થયેલા GuardResult objects ને aggregate_confidence ને પસાર કરવામાં આવે છે. દરેક પરિણામના confidence score ને તેના guard ના weight વડે ગુણવામાં આવે છે — GuardConfig માં સેટ કરાયેલી એક કિંમત જે calibrated વિશ્વાસને encode કરે છે. સારી રીતે ચકાસાયેલ pattern guard નું weight 0.3 હોઈ શકે, calibrated LLM-as-judge નું 0.5, અને નવા પ્રાયોગિક detector નું 0.2. એકત્રિત weighted sum ને કુલ weight વડે ભાગવાથી, કેટલા guards પૂર્ણ થયા તેને ધ્યાનમાં લીધા વિના, 0.0 અને 1.0 વચ્ચેનો normalized score મળે છે.
પછી તે score ની તુલના chain_threshold સાથે કરવામાં આવે છે. Latency budget આ પગલા સાથે એક મહત્વપૂર્ણ રીતે સંકળાય છે: જો તમામ guards ચાલે તે પહેલાં budget સમાપ્ત થઈ જાય, તો aggregation ફક્ત પૂર્ણ થયેલા guards પર જ આગળ વધે છે. આનો અર્થ એ છે કે chain_threshold ને આંશિક-પરિણામના scenarios ને ધ્યાનમાં રાખીને પસંદ કરવો જોઈએ — એક રૂઢિચુસ્ત threshold એવા કિસ્સાઓની ભરપાઈ કરે છે જ્યાં latency ની જરૂરિયાતો પૂરી કરવા માટે ધીમા, ઊંચા-signal વાળા guards છોડી દેવાયા હતા.
કોડ વૉકથ્રુ
હવે જ્યારે તમે સમજો છો કે layered સંરક્ષણ કોઈપણ એક guard કરતાં શા માટે સારું પ્રદર્શન કરે છે, ત્યારે અહીં જુઓ કે orchestrator અને તેનું confidence aggregation કોડમાં કેવી રીતે લાગુ કરવામાં આવે છે.
Guard Chain Orchestrator ની ડિઝાઇન
Orchestrator guards ની એક યાદીનું સંચાલન કરે છે, જેમાં દરેક guard પાસે priority, confidence aggregation માટે weight અને short-circuit threshold હોય છે. Guards priority ના ક્રમમાં execute થાય છે. જો કોઈ guard તેના short-circuit threshold થી ઉપરનો confidence score પરત કરે, તો chain પછીના guards ચલાવ્યા વિના તરત જ request ને અવરોધે છે. જો કોઈ guard short-circuit ટ્રિગર ન કરે, તો orchestrator તમામ પરિણામો એકત્રિત કરે છે અને weighted confidence score ની ગણતરી કરે છે.
Code snippetpython
1from __future__ import annotations 2from typing import Optional 3from pydantic import BaseModel, Field 4 5class GuardConfig(BaseModel): 6 guard_id: str 7 guard_type: str # "pattern", "llm_judge", "guardrails", "document_scanner" 8 priority: int = Field(ge=0, description="Lower number = higher priority") 9 weight: float = Field(ge=0.0, le=1.0) 10 short_circuit_threshold: float = Field(ge=0.0, le=1.0) 11 timeout_ms: int = Field(default=1000) 12 enabled: bool = True 13 14class GuardResult(BaseModel): 15 guard_id: str 16 confidence: float 17 triggered: bool 18 latency_ms: float 19 20class GuardChainResult(BaseModel): 21 allowed: bool 22 total_confidence: float 23 guard_results: list[GuardResult] 24 short_circuited: bool = False 25 short_circuit_guard: Optional[str] = None 26 total_latency_ms: float 27 28def aggregate_confidence( 29 guard_results: list[GuardResult], 30 guard_configs: dict[str, GuardConfig], 31) -> float: 32 weighted_sum = 0.0 33 total_weight = 0.0 34 for result in guard_results: 35 config = guard_configs[result.guard_id] 36 weighted_sum += result.confidence * config.weight 37 total_weight += config.weight 38 return weighted_sum / total_weight if total_weight > 0 else 0.0
GuardConfig orchestrator ને guard ને schedule કરવા માટે જરૂરી બધું જ સમાવે છે: તેની execution priority, aggregation દરમિયાન તેના confidence score પર લાગુ થતું weight, તરત જ block ટ્રિગર કરતો short-circuit threshold, અને per-guard timeout જેથી ધીમો detector pipeline ને અનિશ્ચિત સમય માટે અટકાવી ન શકે.
aggregate_confidence weighted averaging નું પગલું લાગુ કરે છે. તે દરેક પૂર્ણ થયેલા GuardResult પર iterate કરે છે, guard_configs dictionary માંથી તે guard નું weight શોધે છે, અને weighted_sum તથા total_weight એકત્રિત કરે છે. અંતે ભાગાકાર કરવાથી 0.0 અને 1.0 વચ્ચેનો normalized score મળે છે. આ જ બાબત Lethal Trifecta ફ્રેમવર્કને ભિન્ન વિશ્વાસ વ્યક્ત કરવા દે છે — સારી રીતે ચકાસાયેલ pattern guard નું weight 0.3 હોઈ શકે, calibrated LLM-as-judge નું 0.5, અને નવા પ્રાયોગિક detector નું 0.2, જેથી chain નો અંતિમ નિર્ણય દરેક layer પરના ટીમના વિશ્વાસને પ્રતિબિંબિત કરે.
Latency Budget નું સંચાલન
Guard chain સંપૂર્ણ મૂલ્યાંકન pipeline માટે કુલ latency budget લાગુ કરે છે. જો બાકી રહેલું budget આગામી guard ચલાવવા માટે અપૂરતું હોય, તો orchestrator તેને છોડી દે છે અને પહેલેથી પૂર્ણ થયેલા guards ના આધારે નિર્ણય લે છે. આ latency-સંવેદનશીલ endpoints પર security pipeline ને bottleneck બનતી અટકાવે છે.
Code snippetpython
1import time 2 3def run_guard_chain( 4 user_input: str, 5 guards: list[tuple[GuardConfig, callable]], 6 chain_threshold: float, 7 budget_ms: float, 8) -> GuardChainResult: 9 guard_configs = {cfg.guard_id: cfg for cfg, _ in guards} 10 results: list[GuardResult] = [] 11 start = time.monotonic() 12 13 for config, evaluate in sorted(guards, key=lambda g: g[0].priority): 14 if not config.enabled: 15 continue 16 elapsed_ms = (time.monotonic() - start) * 1000 17 if elapsed_ms + config.timeout_ms > budget_ms: 18 break # skip remaining guards to respect latency budget 19 20 t0 = time.monotonic() 21 confidence = evaluate(user_input) 22 latency = (time.monotonic() - t0) * 1000 23 24 result = GuardResult( 25 guard_id=config.guard_id, 26 confidence=confidence, 27 triggered=confidence >= config.short_circuit_threshold, 28 latency_ms=latency, 29 ) 30 results.append(result) 31 32 if result.triggered: 33 return GuardChainResult( 34 allowed=False, 35 total_confidence=confidence, 36 guard_results=results, 37 short_circuited=True, 38 short_circuit_guard=config.guard_id, 39 total_latency_ms=(time.monotonic() - start) * 1000, 40 ) 41 42 total_confidence = aggregate_confidence(results, guard_configs) 43 return GuardChainResult( 44 allowed=total_confidence < chain_threshold, 45 total_confidence=total_confidence, 46 guard_results=results, 47 total_latency_ms=(time.monotonic() - start) * 1000, 48 )
Guards ને priority મુજબ sort કરવામાં આવે છે જેથી સૌથી સસ્તા, સૌથી ઝડપી detectors — સામાન્ય રીતે 1 ms થી ઓછા સમયમાં પૂર્ણ થતા pattern matchers — પહેલા ચાલે. ઊંચા-confidence વાળો hit તરત જ short-circuit કરે છે, જેથી LLM-as-judge અથવા semantic scanner layers ની latency સંપૂર્ણપણે ટાળી શકાય. ફક્ત અસ્પષ્ટ inputs જ સંપૂર્ણ chain માંથી આગળ વધે છે અને chain_threshold સામે weighted નિર્ણય માટે aggregate_confidence સુધી પહોંચે છે.
ખાતરી કરો કે જાણીતી injection string સાથે run_guard_chain ને call કરવાથી, જ્યારે pattern guard ટ્રિગર થાય ત્યારે, પરત કરાયેલ GuardChainResult.allowed False અને short_circuited True હોય, અને chain માં વહેલા પકડાયેલા inputs માટે total_latency_ms configured budget_ms કરતાં ઘણું ઓછું રહે.
શું કરવું અને શું ન કરવું
હવે જ્યારે તમે implementation માંથી પસાર થઈ ગયા છો, ત્યારે નીચેની પ્રથાઓ ટકાઉ અભિગમને નાજુક અભિગમથી અલગ પાડે છે.
શું કરવું
- Guards ને priority કિંમતો સોંપો જેથી સૌથી ઝડપી detectors પહેલા ચાલે — 1 ms થી ઓછા સમયમાં પૂર્ણ થતા pattern matchers ને
GuardConfigમાં સૌથી ઓછોpriorityinteger આપવો જોઈએ, જેથીrun_guard_chainધીમા LLM-as-judge અથવા semantic scanner ને બોલાવતા પહેલાં જ short-circuit નિર્ણય પર પહોંચે, અને મોટાભાગના દૂષિત inputs માટે કુલ latencybudget_msકરતાં ઘણી ઓછી રહે. - દરેક guard ના
GuardConfigમાંનાweightને તમારી ટીમના તે detector પરના calibrated વિશ્વાસને પ્રતિબિંબિત કરવા માટે tune કરો —aggregate_confidenceweighted average ઉત્પન્ન કરે છે, તેથીweight=0.5પર સારી રીતે validated LLM-as-judgeweight=0.2પરના નવા પ્રાયોગિક detector પર યોગ્ય રીતે પ્રભુત્વ ધરાવે છે, અનેchain_thresholdસામે chain નો અંતિમ નિર્ણય નિષ્કપટ બહુમતી મત ને બદલે વાસ્તવિક ભિન્ન વિશ્વાસનું પ્રતિનિધિત્વ કરે છે. - દરેક guard માટે
timeout_msસેટ કરો અનેrun_guard_chainમાં કુલbudget_msલાગુ કરો — orchestrator એવા કોઈપણ guard ને છોડી દે છે જેનોtimeout_mselapsed_msનેbudget_msથી આગળ ધકેલી દે, જેથી LLM-as-judge call ધીમો હોય ત્યારે પણ security pipeline પ્રોડક્શન endpoints પર ક્યારેય latency bottleneck ન બને.
શું ન કરવું
- એક જ guard પર આધાર રાખીને તેને defense-in-depth ન કહો — જો pattern matcher તમારું એકમાત્ર layer હોય, તો એક adversarial bypass (નવું encoding અથવા retrieved document દ્વારા indirect injection) આગળ કોઈ તપાસ વિના
allowed=Trueપરત કરે છે; Lethal Trifecta chain ચોક્કસ એટલા માટે જ અસ્તિત્વમાં છે કારણ કે દરેક guard પ્રકારની અલગ અંધ જગ્યાઓ હોય છે જેની ભરપાઈ બીજા કરે છે. - જ્યારે તમારા detectors ની precision અર્થપૂર્ણ રીતે અલગ હોય ત્યારે તમામ guard ના
weightની કિંમતો સરખી ન રાખો — સરખા weightsaggregate_confidenceને સપાટ average આપે છે, જે તમારા સૌથી વિશ્વસનીય classifier ના signal ને ભૂંસી નાખે છે અને ઊંચા-confidence વાળા LLM-as-judge પરિણામને નબળી રીતે calibrated પ્રાયોગિક detector દ્વારાchain_thresholdથી નીચે પાતળું થવા દે છે. - Priority loop ની અંદર
short_circuit_thresholdની તપાસ છોડી ન દો —run_guard_chainમાં વહેલી-returnશાખાને છોડી દેવાથી દરેક input, સ્પષ્ટ injections સહિત, સંપૂર્ણ chain માંથી અનેaggregate_confidenceમાં જવા મજબૂર થાય છે, જે બિનજરૂરી latency ઉમેરે છે અને પછીના guards ને weighted score નેchain_thresholdથી નીચે લાવવાની તક આપે છે, જ્યારે પહેલા guard પાસે block કરવા માટે પહેલેથી જ નિર્ણાયક પુરાવા હતા.
3 hands-on labs come with this lesson — real code, in a cloud IDE. Create a free account to run them. No card.
Free account · no card · straight to the labs
Or get the full path — from
Listen to this lesson
Audio overviews of this lesson's labs and its chapter, from GenBodha Bytes.
- Guard Chain OrchestratorLab5 min
- Short-Circuit Logic for Guard ChainLab5 min
- Guard Result Aggregator with Weighted ScoringLab5 min
- Prompt Injection DefenseChapter overview20 min
More free lessons in AI Security Engineering
- Ch 1Build prompt injection classifier using LLM-as-judge via LiteLLM
- Ch 1Implement input sanitization pipeline with NeMo Guardrails
- Ch 1Detect indirect injection in RAG-retrieved documents
- Ch 1Build defense-in-depth with layered guard chainYou are here
- Ch 1Deploy injection defense as FastAPI sidecar on GKE
- Ch 1Monitor injection attempts with Prometheus and Grafana
- Ch 3Deploy output sanitizer as response middleware on GKE