Free lesson · GenAI Agent Engineering
Gemini API ને કૉલ કરીને structured responses પરત કરતી Python app લખો
platform proxy દ્વારા Gemini API ને prompts મોકલતી અને response ને format કરતી Python CLI application બનાવો. આ એ જ app છે જેને તમે આખા chapter દરમિયાન containerize કરશો.
Course: Kubernetes Essentials for GenAI Engineers · Chapter 1 · Containerizing LLM Applications
Free to read — no subscription required.
પરિચય
જ્યારે તમે પહેલી વાર Python સર્વિસને Gemini API સાથે જોડો છો, ત્યારે લાલચ એ હોય છે કે API key ને કોઈ constant માં મૂકી દો, તમારા route handler માં request પેસ્ટ કરી દો, અને શિપ કરી દો. એ એક વાર કામ કરે છે — જ્યાં સુધી તમારે તેને deploy કરવાની જરૂર ન પડે. તમે containerize કરો તે ક્ષણે hardcoded secrets તૂટી જાય છે, ad-hoc error handling ને કારણે Gemini ની દરેક નાની અડચણ 500 જેવી દેખાય છે, અને per-request client init દરેક કૉલ પર latency બગાડે છે. જે ટીમો આ શિસ્ત છોડી દે છે તેઓ key rotate કરવા માટે image ફરીથી બનાવતી રહે છે અને એવી તૂટક-તૂટક નિષ્ફળતાઓનો પીછો કરતી રહે છે જેનો કોઈ સુસંગત error shape નથી હોતો.
આ પાઠના અંત સુધીમાં તમે એવી Python એપ્લિકેશન લખી શકશો જે પોતાનું Gemini configuration environment variables માંથી લોડ કરે, FastAPI ના lifespan hook દ્વારા client ને બરાબર એક જ વાર initialize કરે, Gemini API ને prompts મોકલે, અને અનુમાનિત error semantics સાથે structured JSON responses પરત કરે — જે આગલા પાઠમાં સીધી container માં મૂકવા માટે તૈયાર હોય.
મુખ્ય પરિભાષા
- Gemini API — Gemini પરિવારના large language models માટે Google નું HTTPS endpoint. આ પાઠની એપ્લિકેશન તેને prompt મોકલે છે અને જનરેટ થયેલો ટેક્સ્ટ પરત કરે છે; તેના request/response shape ને જાણવાથી જ નીચેનો wrapper કોડ સમજાય છે.
- google-generativeai — અધિકૃત Python SDK જે Gemini HTTP API ને wrap કરે છે, અને prompt-in / text-out કૉલ્સ માટે
GenerativeModel.generate_contentપૂરું પાડે છે. નીચેનોGeminiClientclass તેની ઉપરનું એક પાતળું layer છે. - FastAPI — અહીં વપરાયેલું async Python web framework, જે
/generateendpoint ઉજાગર કરે છે જેથી બાહ્ય callers સીધા Gemini SDK સાથે વાત કર્યા વિના HTTP પર prompts મોકલી શકે. - Pydantic BaseSettings — configuration loader જે Gemini API key, model name અને timeout ને environment variables માંથી વાંચે છે અને જો કોઈ જરૂરી મૂલ્ય ખૂટતું હોય તો startup વખતે જ તરત નિષ્ફળ જાય છે.
- lifespan context — FastAPI નું hook જે એપના જીવનકાળની આસપાસ setup અને teardown ચલાવે છે. અહીં તેનો ઉપયોગ દરેક request પર નહીં પણ startup વખતે એક જ વાર Gemini client બનાવવા માટે થાય છે.
સંકલ્પનાઓ
Layered Architecture
સર્વિસ ત્રણ layers માં વહેંચાય છે, દરેકની પોતાની configuration ચિંતા છે. FastAPI HTTP અને Pydantic validation સંભાળે છે; GeminiClient class API communication અને error translation ની માલિકી ધરાવે છે; Settings module environment variables માંથી configuration લોડ કરે છે. આ વિભાજન secret (API key) ને request-handling logic થી અલગ રાખે છે, અને કોડને સ્પર્શ્યા વિના call shape (timeout, max tokens) ને tune કરવા દે છે.
જો validation input ને નકારે — prompt ખૂટતું હોય, ખાલી string હોય, prompt મર્યાદા કરતાં મોટું હોય — તો FastAPI કોઈ પણ Gemini કૉલ થાય તે પહેલાં 422 પરત કરે છે, જે quota વપરાશને અનુમાનિત રાખે છે.
Environment Variables દ્વારા Configuration
pydantic_settings માંથી BaseSettings દરેક config field ને type, default અને numeric range validators સાથે જાહેર કરે છે. Startup વખતે તે environment variables માંથી વાસ્તવિક મૂલ્યો વાંચે છે, અને development માં .env ફાઇલ પર fallback કરે છે. જરૂરી fields પોતાના default તરીકે ... વાપરે છે, તેથી GEMINI_API_KEY ખૂટતી હોય તો પહેલી request પર અસ્પષ્ટ runtime error આપવાને બદલે startup તરત જ નિષ્ફળ જાય છે. (જુઓ Code Walkthrough.)
Lifespan-Managed Client Initialization
FastAPI નું lifespan async context manager એપના જીવનકાળની આસપાસ એક જ વાર ચાલે છે. ઉદાહરણ તેનો ઉપયોગ એક GeminiClient બનાવવા અને તેને app.state પર મૂકવા માટે કરે છે. દરેક request એ client ને — અને SDK ના અંતર્ગત connection pool ને — દરેક કૉલ પર ફરીથી initialize કરવાને બદલે ફરીથી વાપરે છે. આનો અર્થ એ પણ છે કે credential નિષ્ફળતાઓ user traffic દરમિયાન નહીં પણ startup વખતે જ સામે આવે છે.
Boundary પર Error Translation
Gemini SDK network timeouts, auth failures અને rate limits પર exception ફેંકી શકે છે. generate method કૉલને try/except માં wrap કરે છે અને HTTPException(502) તરીકે re-raise કરે છે, જેથી SDK ની દરેક failure mode એક જ અનુમાનિત HTTP response પર map થાય છે. એક અલગ guard એ કેસ સંભાળે છે જ્યાં Gemini કોઈ parts વગરનો response પરત કરે (સામાન્ય રીતે safety-filtered output) અને SDK ના અસ્પષ્ટ ValueError ને બદલે સ્પષ્ટ સંદેશ સાથે એ જ 502 આપે છે.
Code Walkthrough
આ walkthrough ચારેય સંકલ્પનાઓને સાથે દર્શાવે છે: BaseSettings દ્વારા environment-driven configuration, lifespan-bootstrapped client initialization, request validation, અને boundary પર error translation. પહેલો snippet Settings જાહેર કરે છે; બીજો તેને GeminiClient સાથે FastAPI એપમાં જોડે છે.
Code snippetpython
1# settings.py 2from functools import lru_cache 3 4from pydantic import Field 5from pydantic_settings import BaseSettings, SettingsConfigDict 6 7class Settings(BaseSettings): 8 model_config = SettingsConfigDict(env_file=".env", case_sensitive=False) 9 10 gemini_api_key: str = Field(..., description="API key for Gemini") 11 gemini_model: str = Field(default="gemini-1.5-flash") 12 request_timeout: int = Field(default=30, ge=5, le=120) 13 max_output_tokens: int = Field(default=1024, ge=1, le=8192) 14 temperature: float = Field(default=0.7, ge=0.0, le=2.0) 15 16@lru_cache 17def get_settings() -> Settings: 18 return Settings()
gemini_api_key પોતાના default તરીકે ... વાપરે છે, જે તેને જરૂરી બનાવે છે — જો GEMINI_API_KEY સેટ ન હોય તો એપ startup વખતે જ ઝડપથી નિષ્ફળ જાય છે. ge/le validators અમાન્ય numeric મૂલ્યોને Gemini સુધી પહોંચતા પહેલાં રોકે છે. @lru_cache સમગ્ર process માં એક જ shared Settings instance ની ખાતરી કરે છે.
Code snippetpython
1# main.py 2from contextlib import asynccontextmanager 3 4import google.generativeai as genai 5from fastapi import FastAPI, HTTPException 6from pydantic import BaseModel, Field 7 8from settings import get_settings 9 10class PromptRequest(BaseModel): 11 prompt: str = Field(..., min_length=1, max_length=10000) 12 temperature: float | None = Field(default=None, ge=0.0, le=2.0) 13 14class GeminiClient: 15 def __init__(self, settings): 16 genai.configure(api_key=settings.gemini_api_key) 17 self._model = genai.GenerativeModel(settings.gemini_model) 18 self._settings = settings 19 20 def generate(self, prompt: str, temperature: float | None = None) -> dict: 21 config = genai.types.GenerationConfig( 22 max_output_tokens=self._settings.max_output_tokens, 23 temperature=temperature or self._settings.temperature, 24 ) 25 try: 26 response = self._model.generate_content( 27 prompt, 28 generation_config=config, 29 request_options={"timeout": self._settings.request_timeout}, 30 ) 31 except Exception as exc: 32 raise HTTPException(status_code=502, detail=str(exc)) from exc 33 34 if not response.parts: 35 raise HTTPException(status_code=502, detail="Empty response from Gemini") 36 37 return { 38 "model": self._settings.gemini_model, 39 "prompt": prompt, 40 "generated_text": response.text, 41 } 42 43@asynccontextmanager 44async def lifespan(app: FastAPI): 45 app.state.client = GeminiClient(get_settings()) 46 yield 47 48app = FastAPI(title="Gemini LLM Service", lifespan=lifespan) 49 50@app.post("/generate") 51async def generate_text(request: PromptRequest): 52 return app.state.client.generate( 53 prompt=request.prompt, 54 temperature=request.temperature, 55 ) 56 57@app.get("/health") 58async def health_check(): 59 return {"status": "healthy"}
PromptRequest edge પર 1-થી-10,000-અક્ષરોનું prompt લાગુ કરે છે, તેથી ખાલી કે વધુ પડતી મોટી request ક્યારેય Gemini સુધી પહોંચતી નથી. lifespan context startup વખતે એક જ વાર GeminiClient બનાવે છે અને તેને app.state પર સંગ્રહે છે. generate ની અંદર, try/except કોઈ પણ SDK exception ને HTTP 502 માં અનુવાદિત કરે છે, અને response.parts ની તપાસ .text exception ફેંકે તે પહેલાં safety-filtered કેસને પકડી લે છે. /health endpoint Gemini ને કૉલ કર્યા વિના પરત ફરે છે, જે container probes ને સસ્તું લક્ષ્ય આપે છે.
તમને ખબર પડશે કે તે કામ કરે છે જ્યારે, GEMINI_API_KEY export કરીને અને uvicorn main:app --reload ચલાવ્યા પછી, http://localhost:8000/generate પર Content-Type: application/json અને body {"prompt":"Say hello in one short sentence."} સાથેની POST request બિન-ખાલી generated_text field સાથે 200 પરત કરે.
માટે વ્યવહારમાં
ઉપરની pattern — environment-driven config, lifespan-bootstrapped client, edge પર validation, boundary પર error translation — કોઈ પણ Gemini-backed Python સર્વિસ પર લાગુ પડે છે. તમે તેને કેવી રીતે વિસ્તારો છો તે તમે કઈ ભૂમિકા માટે બનાવી રહ્યા છો તેના પર આધાર રાખે છે.
શું કરવું અને શું ન કરવું
શું કરવું
- API key ને environment variable માંથી લોડ કરો — secret ને source control થી બહાર રાખે છે અને image ફરીથી બનાવ્યા વિના દરેક environment માટે keys બદલવા દે છે.
- Gemini client ને lifespan hook માં એક જ વાર initialize કરો — SDK ના connection pool નો ફરીથી ઉપયોગ કરે છે અને credential નિષ્ફળતાઓને traffic ની વચ્ચે નહીં પણ startup વખતે જ સામે લાવે છે.
- SDK exceptions ને HTTP 502 માં અનુવાદિત કરો — અંતર્ગત સમસ્યા timeout, auth failure કે rate limit હોય, callers ને એક જ અનુમાનિત failure mode આપે છે.
શું ન કરવું
- API key કે model name ને hardcode ન કરો — દરેક environment ફેરફાર image rebuild બની જાય છે, અને leak થયેલી image production credentials leak કરે છે.
- પહેલાં
response.partsતપાસ્યા વિનાresponse.textને કૉલ ન કરો — safety-filtered Gemini response એ parts વગરનો માન્ય object છે, અને.textઅસ્પષ્ટValueErrorફેંકે છે. - liveness probes માટે
/generateendpoint નો ફરીથી ઉપયોગ ન કરો — દરેક probe Gemini quota બગાડે છે; એક અલગ/healthroute ઉજાગર કરો જે API ને કૉલ કર્યા વિના પરત ફરે.
A hands-on lab comes with this lesson — real code, in a cloud IDE. Create a free account to run it. No card.
Free account · no card · straight to the lab
Or get the full path — from
Listen to this lesson
Audio overviews of this lesson's labs and its chapter, from GenBodha Bytes.
- Structured Text Report GeneratorLab5 min
- Containerizing LLM ApplicationsChapter overview23 min
More free lessons in Kubernetes Essentials for GenAI Engineers
- Ch 1Write a Python app that calls the Gemini API and returns structured responsesYou are here
- Ch 1Write a Dockerfile and build a container image for the LLM app
- Ch 1Use Docker Compose to run the LLM app with supporting services
- Ch 2Deploy the LLM app as your first Kubernetes pod
- Ch 4Manage deployment lifecycle with kubectl rollout
- Ch 9Create a Helm chart for the LLM chat application
- Ch 9Use Kustomize bases and overlays for the LLM app