GenerationConfig
Optional load-time tuning for an on-device LLM. Every field is nullable — null means "leave the backend's own default alone". Backends apply what their engine exposes and ignore the rest:
MediaPipe (Gemma) — maxTokens + accelerator + the topK ceiling are load-time options; the topK/topP/temperature sampler is applied per inference session.
ML Kit GenAI (Gemini Nano) — topK/temperature/maxTokens are wired per-request (MlKitGenAiOnDeviceLlm's
buildRequest); topP/accelerator have no equivalent on this API and are ignored.Foundation Models (iOS) — FoundationModelsOnDeviceLlm's Swift bridge (
ai/ios-bridge/) callsLanguageModelSessionthrough a completion-handler shape that carries no topK/topP/temperature/maxTokens knobs, so every field is ignored here too. Kept uniform so a caller wires one config regardless of backend.