OnDeviceLlm
A single-shot on-device text LLM tier. Kept deliberately tiny — one availability gate and one text-in/text-out call — so each platform's actual (ML Kit GenAI on Android, Foundation Models on iOS, unavailable elsewhere) is a thin wrapper, and DefaultJobIntelligence never has to know which backend ran.
generate returns a typed AiFailure on any failure — AiFailure.ModelNotResident (download it) is distinguishable from AiFailure.NotSupportedOnPlatform (this device never can) — so the caller can either degrade to its own heuristic tier or say the right thing to the user.
Inheritors
Properties
True when this backend accepts an LlmPart.Image in generate. False = text-only.
Functions
Honest self-report for a capability-aware caller: whether this backend genuinely streams tokens and accepts images, which GenerationConfig fields it reads, and — when it can't run — the real AiFailure reason instead of a bare false. Default answer is conservative (no streaming, no honored config) and derives AiCapabilities.unavailableReason from isAvailable alone; a backend with a more specific gate (model residency, OS feature status) overrides this to report why.
Multimodal entry point. Default maps a single LlmPart.Text onto generate (String) so every existing text-only actual (MediaPipe, Foundation Models, jvm) keeps working with zero changes. Backends that accept images (ML Kit GenAI) override this directly.
Streaming variant of generate. Default replays the single-shot result as one emission so every existing actual keeps working with zero changes; backends with native token streaming (ML Kit GenAI, MediaPipe) override this directly. CompositeOnDeviceLlm overrides both overloads too, so streaming through it reaches whichever backend it picked, not this default.
Cheap, synchronous floor — true only when the platform could run inference. NOT the authoritative runtime check (model residency is async); generate still guards internally.