CompositeOnDeviceLlm
Detection-ordered OnDeviceLlm (ai-engineering.md §7): probes an ordered list of backends and uses the first one that both reports available AND actually produces output. This is how a device escalates ML Kit Gemini Nano (AICore-only) → MediaPipe Gemma (broad device coverage, downloaded on demand) → (nothing → the heuristic tier takes over upstream in DefaultJobIntelligence).
The seam is unchanged: callers still see one OnDeviceLlm; the ordering lives here.
Properties
True when this backend accepts an LlmPart.Image in generate. False = text-only.
Functions
Runs prompt on-device. Returns the model's text, or null when unavailable/declined/failed.
Multimodal entry point. Default maps a single LlmPart.Text onto generate (String) so every existing text-only actual (MediaPipe, Foundation Models, jvm) keeps working with zero changes. Backends that accept images (ML Kit GenAI) override this directly.
Streaming variant of generate. Default replays the single-shot result as one emission so every existing actual keeps working with zero changes; backends with native token streaming (ML Kit GenAI) override this directly.
Cheap, synchronous floor — true only when the platform could run inference. NOT the authoritative runtime check (model residency is async); generate still guards internally.