Free lesson · GenAI Security Engineering
அடுக்கு guard chain மூலம் defense-in-depth உருவாக்குங்கள்
short-circuit logic, confidence aggregation மற்றும் configurable thresholds உடன் பல injection detectors-ஐ ஒரு அடுக்கு pipeline-ஆக ஒருங்கிணைக்கவும்.
Course: AI Security Engineering · Chapter 1 · Prompt Injection Defense
Free to read — no subscription required.
அறிமுகம்
Production-இல் ஒரே ஒரு injection detector-ஐ மட்டும் நம்பியிருக்கும்போது, தவறவிடப்பட்ட ஒரு pattern அல்லது ஒரு adversarial bypass உங்கள் முழு பாதுகாப்பையும் தோற்கடித்துவிடும். Lethal Trifecta framework இதை, பல சுயாதீன guard-களை — pattern matcher-கள், LLM-as-judge classifier-கள், மற்றும் semantic scanner-கள் — ஒரு pipeline-இல் சங்கிலியாக இணைப்பதன் மூலம் தீர்க்கிறது; அங்கு ஒவ்வொரு அடுக்கும் மற்றவற்றின் குருட்டுப் புள்ளிகளை ஈடுசெய்கிறது. இந்தப் பாடத்தின் முடிவில், short-circuit logic, ஒவ்வொரு guard-க்கும் தனித்தனி confidence weighting, மற்றும் latency budget மேலாண்மை கொண்ட ஒரு guard chain orchestrator-ஐ நீங்கள் உருவாக்க முடியும்; இது நேரடி மற்றும் மறைமுக தாக்குதல் வழிகள் இரண்டிலும் prompt injection முயற்சிகளைத் தடுக்கும்.
முக்கிய சொற்கள்
- Lethal Trifecta — மூன்று வகையான சுயாதீன guard-களை — pattern matcher-கள், LLM-as-judge classifier-கள், மற்றும் semantic scanner-கள் — சங்கிலியாக இணைக்கும் ஒரு defense-in-depth framework; இதனால் ஒவ்வொரு அடுக்கின் கவரேஜும் மற்றவற்றின் குருட்டுப் புள்ளிகளை ஈடுசெய்கிறது.
- Guard Priority —
GuardConfig-இல் உள்ள ஒரு எண் புலம் (குறைந்த மதிப்பு = அதிக முன்னுரிமை); இது chain-க்குள் செயல்படுத்தும் வரிசையைத் தீர்மானிக்கிறது, மலிவான வேகமான detector-கள் LLM-as-judge போன்ற latency-அதிக detector-களுக்கு முன் இயங்குவதை உறுதி செய்கிறது. - Short-Circuit Threshold —
GuardConfig-இல் உள்ள ஒவ்வொரு guard-க்கானshort_circuit_thresholdமதிப்பு; ஒரு guard திருப்பித் தரும் confidence இதைச் சமமாகவோ அல்லது மீறவோ செய்தால்,run_guard_chainமீதமுள்ள எந்த guard-ஐயும் அழைக்காமல் உடனடியாகallowed=Falseஎன்று திருப்பித் தருகிறது. - Latency Budget —
run_guard_chain-க்கு அனுப்பப்படும்budget_msஉச்சவரம்பு; கழிந்த நேரமும் ஒரு guard-இன்timeout_ms-உம் சேர்ந்து இந்த மதிப்பை மீறினால், orchestrator அந்த guard-ஐத் தவிர்த்துவிட்டு, இதுவரை சேகரிக்கப்பட்ட முடிவுகளின் மீது aggregation-க்குச் செல்கிறது. - Weighted Confidence Aggregation — நிறைவடைந்த அனைத்து
GuardResultscore-களின் மீதும்aggregate_confidenceகணக்கிடும் normalized weighted average; இதில் ஒவ்வொரு guard-இன் பங்களிப்பும்GuardConfig-இல் உள்ள அதன்weight-ஆல் அளவிடப்படுகிறது, இது guard வகைகளுக்கிடையே வேறுபட்ட நம்பிக்கையை வெளிப்படுத்துகிறது. - Chain Threshold — இறுதி aggregated confidence score-உடன் ஒப்பிடப்படும்
chain_thresholdparameter; weighted score இதற்குக் கீழே இருக்கும் input-கள் அனுமதிக்கப்படுகின்றன, இதற்குச் சமமாகவோ மேலேயோ இருப்பவை தடுக்கப்படுகின்றன.
கருத்துகள்
ஒற்றை Guard-ஐ விட சுயாதீன அடுக்குகள் ஏன் சிறப்பாகச் செயல்படுகின்றன
Production-க்கு எந்த ஒரு injection detector-உம் தனியாக போதுமான நம்பகத்தன்மை கொண்டதல்ல. Pattern matcher-கள் வேகமானவை, ஆனால் அறியப்பட்ட signature-களுடன் பொருந்தாத புதிய அல்லது மறைக்கப்பட்ட தாக்குதல்களைத் தவறவிடுகின்றன. LLM-as-judge classifier-கள் சிறப்பாக generalize செய்கின்றன, ஆனால் பத்து முதல் நூறு மில்லிவினாடிகள் வரை latency-ஐச் சேர்க்கின்றன; மேலும் ஒரு மொழி மாதிரியைக் குழப்புவதற்காக வடிவமைக்கப்பட்ட adversarial input-களால் அவையே தவறாக வழிநடத்தப்படலாம். Semantic scanner-கள் pattern rule-களால் ஒருபோதும் வெளிப்படுத்த முடியாத embedding-space ஒழுங்கின்மைகளைப் பிடிக்கின்றன, ஆனால் அவற்றின் துல்லியம் reference corpus உங்கள் threat surface-ஐ எவ்வளவு நன்றாகக் கவர் செய்கிறது என்பதைப் பொறுத்தது. ஒவ்வொரு detector-க்கும் ஒரு தனித்துவமான தோல்வி முறை உள்ளது — உங்கள் system-ஐ ஆய்வு செய்யும் ஒரு தாக்குபவர், அதுவே உங்கள் ஒரே அடுக்கு என்றால், மிகவும் பலவீனமான அடுக்கை மட்டும் தோற்கடித்தால் போதும்.
Lethal Trifecta framework சுயாதீனமான தோல்வி முறைகளை பாதுகாப்பின் கட்டமைப்பு ரீதியான பலமாக மாற்றுகிறது. வெவ்வேறு methodology கொண்ட மூன்று detector-கள் ஒவ்வொன்றும் ஒரு தாக்குதலைத் தவறவிடுவதற்கு சில நிகழ்தகவு இருக்கும்போது, மூன்றும் ஒரே தாக்குதலைத் தவறவிடுவதற்கான நிகழ்தகவு அவற்றின் தனிப்பட்ட miss rate-களின் பெருக்கல் ஆகும் — எந்த ஒரு detector-ஐ விடவும் மிகக் குறைவு. Guard chain architecture (Code Walkthrough பார்க்கவும்) இந்த சுயாதீன guard-களை ஒன்றாக இணைப்பதற்கான orchestration machinery-ஐ வழங்குகிறது; அதே நேரத்தில் latency-ஐ முன்கணிக்கக்கூடியதாகவும், ஒவ்வொரு guard-இன் தீர்ப்பையும் குழுவின் calibrated நம்பிக்கையால் weight செய்யப்பட்டதாகவும் வைத்திருக்கிறது.
Guard வரிசைப்படுத்தல் மற்றும் Short-Circuit Logic
மதிப்பீடு தொடங்கும் முன் orchestrator guard-களை அவற்றின் priority புலத்தின் அடிப்படையில் வரிசைப்படுத்துகிறது. வடிவமைப்புக் கொள்கை செலவு சமச்சீரற்ற தன்மை ஆகும்: 1 ms-க்கும் குறைவான நேரத்தில் முடியும் ஒரு pattern matcher, ஒரு LLM call-இன் செலவில் மிகச் சிறிய பகுதியிலேயே அறியப்பட்ட ஒரு injection string-ஐத் திட்டவட்டமாகத் தடுக்க முடியும். மலிவான, மிகத் தீர்க்கமான guard-களை முதலில் வைப்பதன் மூலம், தெளிவான தாக்குதல்கள் விலையுயர்ந்த downstream அடுக்குகளை அடைவதற்கு முன்பே நிறுத்தப்படுகின்றன.
Short-circuit logic இதை முறைப்படுத்துகிறது. ஒவ்வொரு GuardConfig-உம் அதன் சொந்த short_circuit_threshold-ஐக் கொண்டுள்ளது. ஒரு guard அந்த மதிப்புக்குச் சமமான அல்லது அதற்கு மேலான confidence score-ஐத் திருப்பித் தந்த கணமே, run_guard_chain short_circuited=True என அமைத்து allowed=False என்று திருப்பித் தருகிறது — வரிசைப்படுத்தப்பட்ட பட்டியலில் மீதமுள்ள guard-கள் ஒருபோதும் செயல்படுத்தப்படுவதில்லை. வேகமான detector-களைக் குறைந்த confidence-உடன் கடக்கும் input-கள் மட்டுமே முழு chain வழியாக aggregation-ஐ அடைகின்றன.
Confidence Weighting மற்றும் இறுதி முடிவு
எந்த guard-உம் short-circuit-ஐத் தூண்டாதபோது, நிறைவடைந்த அனைத்து GuardResult object-களும் aggregate_confidence-க்கு அனுப்பப்படுகின்றன. ஒவ்வொரு result-இன் confidence score-உம் அதன் guard-இன் weight-ஆல் பெருக்கப்படுகிறது — இது GuardConfig-இல் அமைக்கப்பட்ட, calibrated நம்பிக்கையைக் குறிக்கும் ஒரு மதிப்பு. முழுமையாகச் சோதிக்கப்பட்ட ஒரு pattern guard 0.3 weight-ஐயும், calibrated LLM-as-judge 0.5-ஐயும், புதிய பரிசோதனை detector 0.2-ஐயும் கொண்டிருக்கலாம். திரட்டப்பட்ட weighted sum-ஐ மொத்த weight-ஆல் வகுப்பது, எத்தனை guard-கள் நிறைவடைந்தாலும் 0.0 முதல் 1.0 வரையிலான ஒரு normalized score-ஐத் தருகிறது.
பின்னர் அந்த score chain_threshold-உடன் ஒப்பிடப்படுகிறது. Latency budget இந்தப் படியுடன் ஒரு முக்கியமான வகையில் தொடர்பு கொள்கிறது: அனைத்து guard-களும் இயங்கும் முன் budget தீர்ந்துவிட்டால், முடிந்த guard-களின் மீது மட்டுமே aggregation நடக்கிறது. இதன் பொருள், chain_threshold-ஐ பகுதி-முடிவு சூழ்நிலைகளை மனதில் கொண்டு தேர்ந்தெடுக்க வேண்டும் — latency தேவைகளைப் பூர்த்தி செய்ய மெதுவான, அதிக-signal கொண்ட guard-கள் தவிர்க்கப்பட்ட நிகழ்வுகளை ஒரு பழமைவாத threshold ஈடுசெய்கிறது.
Code Walkthrough
அடுக்கு பாதுகாப்பு எந்த ஒற்றை guard-ஐ விடவும் ஏன் சிறப்பாகச் செயல்படுகிறது என்பதை இப்போது நீங்கள் புரிந்துகொண்டீர்கள்; orchestrator-உம் அதன் confidence aggregation-உம் code-இல் எப்படிச் செயல்படுத்தப்படுகின்றன என்பது இதோ.
Guard Chain Orchestrator வடிவமைப்பு
Orchestrator guard-களின் ஒரு பட்டியலை நிர்வகிக்கிறது; ஒவ்வொன்றும் ஒரு priority, confidence aggregation-க்கான ஒரு weight, மற்றும் ஒரு short-circuit threshold-ஐக் கொண்டுள்ளது. Guard-கள் priority வரிசையில் செயல்படுத்தப்படுகின்றன. ஒரு guard அதன் short-circuit threshold-ஐ விட அதிகமான confidence score-ஐத் திருப்பித் தந்தால், chain அடுத்தடுத்த guard-களை இயக்காமல் உடனடியாக request-ஐத் தடுக்கிறது. எந்த guard-உம் short-circuit-ஐத் தூண்டாவிட்டால், orchestrator அனைத்து முடிவுகளையும் சேகரித்து ஒரு weighted confidence score-ஐக் கணக்கிடுகிறது.
Code snippetpython
1from __future__ import annotations 2from typing import Optional 3from pydantic import BaseModel, Field 4 5class GuardConfig(BaseModel): 6 guard_id: str 7 guard_type: str # "pattern", "llm_judge", "guardrails", "document_scanner" 8 priority: int = Field(ge=0, description="Lower number = higher priority") 9 weight: float = Field(ge=0.0, le=1.0) 10 short_circuit_threshold: float = Field(ge=0.0, le=1.0) 11 timeout_ms: int = Field(default=1000) 12 enabled: bool = True 13 14class GuardResult(BaseModel): 15 guard_id: str 16 confidence: float 17 triggered: bool 18 latency_ms: float 19 20class GuardChainResult(BaseModel): 21 allowed: bool 22 total_confidence: float 23 guard_results: list[GuardResult] 24 short_circuited: bool = False 25 short_circuit_guard: Optional[str] = None 26 total_latency_ms: float 27 28def aggregate_confidence( 29 guard_results: list[GuardResult], 30 guard_configs: dict[str, GuardConfig], 31) -> float: 32 weighted_sum = 0.0 33 total_weight = 0.0 34 for result in guard_results: 35 config = guard_configs[result.guard_id] 36 weighted_sum += result.confidence * config.weight 37 total_weight += config.weight 38 return weighted_sum / total_weight if total_weight > 0 else 0.0
GuardConfig ஒரு guard-ஐ schedule செய்ய orchestrator-க்குத் தேவையான அனைத்தையும் பிடிக்கிறது: அதன் execution priority, aggregation-இன் போது அதன் confidence score-க்குப் பயன்படுத்தப்படும் weight, உடனடித் தடையைத் தூண்டும் short-circuit threshold, மற்றும் ஒரு மெதுவான detector pipeline-ஐ காலவரையின்றி நிறுத்திவிடாதபடி ஒவ்வொரு guard-க்கும் ஒரு timeout.
aggregate_confidence weighted averaging படியைச் செயல்படுத்துகிறது. இது நிறைவடைந்த ஒவ்வொரு GuardResult-இன் மீதும் iterate செய்து, guard_configs dictionary-இலிருந்து அந்த guard-இன் weight-ஐத் தேடி, weighted_sum மற்றும் total_weight-ஐத் திரட்டுகிறது. இறுதியில் வகுப்பது 0.0 முதல் 1.0 வரையிலான ஒரு normalized score-ஐத் தருகிறது. இதுவே Lethal Trifecta framework வேறுபட்ட நம்பிக்கையை வெளிப்படுத்த அனுமதிக்கிறது — முழுமையாகச் சோதிக்கப்பட்ட ஒரு pattern guard 0.3 weight-ஐயும், calibrated LLM-as-judge 0.5-ஐயும், புதிய பரிசோதனை detector 0.2-ஐயும் கொண்டிருக்கலாம்; இதனால் chain-இன் இறுதி முடிவு ஒவ்வொரு அடுக்கிலும் குழுவின் நம்பிக்கையைப் பிரதிபலிக்கிறது.
Latency Budget மேலாண்மை
Guard chain முழு மதிப்பீட்டு pipeline-க்கும் ஒரு மொத்த latency budget-ஐ அமல்படுத்துகிறது. அடுத்த guard-ஐ இயக்க மீதமுள்ள budget போதுமானதாக இல்லாவிட்டால், orchestrator அதைத் தவிர்த்துவிட்டு, ஏற்கனவே நிறைவடைந்த guard-களின் அடிப்படையில் ஒரு முடிவை எடுக்கிறது. இது latency-உணர்திறன் கொண்ட endpoint-களில் security pipeline ஒரு bottleneck ஆவதைத் தடுக்கிறது.
Code snippetpython
1import time 2 3def run_guard_chain( 4 user_input: str, 5 guards: list[tuple[GuardConfig, callable]], 6 chain_threshold: float, 7 budget_ms: float, 8) -> GuardChainResult: 9 guard_configs = {cfg.guard_id: cfg for cfg, _ in guards} 10 results: list[GuardResult] = [] 11 start = time.monotonic() 12 13 for config, evaluate in sorted(guards, key=lambda g: g[0].priority): 14 if not config.enabled: 15 continue 16 elapsed_ms = (time.monotonic() - start) * 1000 17 if elapsed_ms + config.timeout_ms > budget_ms: 18 break # skip remaining guards to respect latency budget 19 20 t0 = time.monotonic() 21 confidence = evaluate(user_input) 22 latency = (time.monotonic() - t0) * 1000 23 24 result = GuardResult( 25 guard_id=config.guard_id, 26 confidence=confidence, 27 triggered=confidence >= config.short_circuit_threshold, 28 latency_ms=latency, 29 ) 30 results.append(result) 31 32 if result.triggered: 33 return GuardChainResult( 34 allowed=False, 35 total_confidence=confidence, 36 guard_results=results, 37 short_circuited=True, 38 short_circuit_guard=config.guard_id, 39 total_latency_ms=(time.monotonic() - start) * 1000, 40 ) 41 42 total_confidence = aggregate_confidence(results, guard_configs) 43 return GuardChainResult( 44 allowed=total_confidence < chain_threshold, 45 total_confidence=total_confidence, 46 guard_results=results, 47 total_latency_ms=(time.monotonic() - start) * 1000, 48 )
Guard-கள் priority-ஆல் வரிசைப்படுத்தப்படுகின்றன, இதனால் மலிவான, வேகமான detector-கள் — பொதுவாக 1 ms-க்கும் குறைவான நேரத்தில் முடியும் pattern matcher-கள் — முதலில் இயங்குகின்றன. அதிக-confidence hit உடனடியாக short-circuit செய்து, LLM-as-judge அல்லது semantic scanner அடுக்குகளின் latency-ஐ முழுமையாகத் தவிர்க்கிறது. தெளிவற்ற input-கள் மட்டுமே முழு chain வழியாகத் தொடர்ந்து, chain_threshold-க்கு எதிரான ஒரு weighted முடிவுக்காக aggregate_confidence-ஐ அடைகின்றன.
அறியப்பட்ட ஒரு injection string-உடன் run_guard_chain-ஐ அழைப்பது, pattern guard தூண்டப்படும்போது திருப்பித் தரப்படும் GuardChainResult.allowed False ஆகவும் short_circuited True ஆகவும் இருப்பதையும், chain-இன் தொடக்கத்திலேயே பிடிக்கப்படும் input-களுக்கு total_latency_ms கட்டமைக்கப்பட்ட budget_ms-க்கு மிகக் குறைவாகவே இருப்பதையும் உறுதிப்படுத்துங்கள்.
செய்ய வேண்டியவை மற்றும் செய்யக்கூடாதவை
நீங்கள் இப்போது implementation-ஐ முழுமையாகக் கடந்துவிட்டீர்கள்; கீழே உள்ள நடைமுறைகள் ஒரு நீடித்த அணுகுமுறையை ஒரு உடையக்கூடிய அணுகுமுறையிலிருந்து பிரிக்கின்றன.
செய்ய வேண்டியவை
- வேகமான detector-கள் முதலில் இயங்கும்படி guard-களுக்கு priority மதிப்புகளை ஒதுக்குங்கள் — 1 ms-க்கும் குறைவான நேரத்தில் முடியும் pattern matcher-கள்
GuardConfig-இல் மிகக் குறைந்தpriorityinteger-ஐக் கொண்டிருக்க வேண்டும்; இதனால்run_guard_chainமெதுவான LLM-as-judge அல்லது semantic scanner-ஐ அழைப்பதற்கு முன்பே ஒரு short-circuit முடிவை அடைகிறது, பெரும்பாலான தீங்கிழைக்கும் input-களுக்கு மொத்த latency-ஐbudget_ms-க்கு மிகக் குறைவாக வைத்திருக்கிறது. - அந்த detector மீதான உங்கள் குழுவின் calibrated நம்பிக்கையைப் பிரதிபலிக்கும்படி
GuardConfig-இல் ஒவ்வொரு guard-இன்weight-ஐயும் tune செய்யுங்கள் —aggregate_confidenceஒரு weighted average-ஐத் தருகிறது; எனவே முழுமையாகச் சரிபார்க்கப்பட்ட ஒரு LLM-as-judgeweight=0.5-இல்,weight=0.2-இல் உள்ள புதிய பரிசோதனை detector-ஐச் சரியாக ஆதிக்கம் செய்கிறது, மேலும்chain_threshold-க்கு எதிரான chain-இன் இறுதி முடிவு ஒரு எளிய majority vote அல்ல, உண்மையான வேறுபட்ட confidence-ஐப் பிரதிபலிக்கிறது. - ஒவ்வொரு guard-க்கும் ஒரு
timeout_msஅமைத்து,run_guard_chain-இல் மொத்தbudget_ms-ஐ அமல்படுத்துங்கள் — எந்த guard-இன்timeout_mselapsed_ms-ஐbudget_ms-ஐத் தாண்டச் செய்யுமோ அந்த guard-ஐ orchestrator தவிர்க்கிறது; இதனால் ஒரு LLM-as-judge call மெதுவாக இருந்தாலும் production endpoint-களில் security pipeline ஒருபோதும் latency bottleneck ஆவதில்லை.
செய்யக்கூடாதவை
- ஒரே ஒரு guard-ஐ நம்பிவிட்டு அதை defense-in-depth என்று அழைக்காதீர்கள் — ஒரு pattern matcher மட்டுமே உங்கள் ஒரே அடுக்கு என்றால், ஒரு adversarial bypass (புதிய encoding அல்லது retrieve செய்யப்பட்ட document வழியாக மறைமுக injection) மேலும் எந்தச் சரிபார்ப்பும் இல்லாமல்
allowed=True-ஐத் திருப்பித் தருகிறது; ஒவ்வொரு guard வகைக்கும் மற்றவை ஈடுசெய்யும் தனித்துவமான குருட்டுப் புள்ளிகள் இருப்பதாலேயே Lethal Trifecta chain உள்ளது. - உங்கள் detector-கள் குறிப்பிடத்தக்க அளவில் வேறுபட்ட precision கொண்டிருக்கும்போது அனைத்து guard
weightமதிப்புகளையும் சமமாக அமைக்காதீர்கள் — சம weight-கள்aggregate_confidence-க்கு ஒரு தட்டையான average-ஐத் தருகின்றன, உங்கள் மிக நம்பகமான classifier-இன் signal-ஐ அழித்து, அதிக-confidence LLM-as-judge முடிவை மோசமாக calibrate செய்யப்பட்ட பரிசோதனை detector-ஆல்chain_threshold-க்குக் கீழே நீர்த்துப்போகச் செய்கின்றன. - Priority loop-க்குள்
short_circuit_thresholdசரிபார்ப்பை விட்டுவிடாதீர்கள் —run_guard_chain-இல் முன்கூட்டிய-returnகிளையைத் தவிர்ப்பது, வெளிப்படையான injection-கள் உட்பட ஒவ்வொரு input-ஐயும் முழு chain வழியாகவும்aggregate_confidence-க்குள்ளும் கட்டாயப்படுத்துகிறது; இது தேவையற்ற latency-ஐச் சேர்க்கிறது, மேலும் முதல் guard-க்கு ஏற்கனவே தடுக்கத் தீர்க்கமான ஆதாரம் இருந்தபோதும், அடுத்தடுத்த guard-கள் weighted score-ஐchain_threshold-க்குக் கீழே குறைக்க ஒரு வாய்ப்பை அளிக்கிறது.
3 hands-on labs come with this lesson — real code, in a cloud IDE. Create a free account to run them. No card.
Free account · no card · straight to the labs
Or get the full path — from
Listen to this lesson
Audio overviews of this lesson's labs and its chapter, from GenBodha Bytes.
- Guard Chain OrchestratorLab5 min
- Short-Circuit Logic for Guard ChainLab5 min
- Guard Result Aggregator with Weighted ScoringLab5 min
- Prompt Injection DefenseChapter overview20 min
More free lessons in AI Security Engineering
- Ch 1Build prompt injection classifier using LLM-as-judge via LiteLLM
- Ch 1Implement input sanitization pipeline with NeMo Guardrails
- Ch 1Detect indirect injection in RAG-retrieved documents
- Ch 1Build defense-in-depth with layered guard chainYou are here
- Ch 1Deploy injection defense as FastAPI sidecar on GKE
- Ch 1Monitor injection attempts with Prometheus and Grafana
- Ch 3Deploy output sanitizer as response middleware on GKE