CompositeOnDeviceLlm
Detection-ordered OnDeviceLlm: probes an ordered list of backends and uses the first one that both reports available AND actually produces output. This is how a device escalates ML Kit Gemini Nano (AICore-only) → MediaPipe Gemma (broad device coverage, downloaded on demand) → CloudOnDeviceLlm (network fallback, appended by the caller only when a key exists) → (nothing → the caller's own heuristic fallback tier).
The seam is unchanged: callers still see one OnDeviceLlm. Every prompt/text part is run through PromptGuard here, before any backend sees it — a caller like JobSummarizer in the README builds one flat string with its own instructions and the untrusted JD/receipt/message concatenated together, with no way for this class to tell which is which; PromptGuard.wrap still delimits the whole thing and reasserts that embedded instruction-shaped text isn't one, and callers cannot bypass this by going through a different overload — every entry point below (single-shot and streaming, text and multimodal) routes through it.
Properties
True when this backend accepts an LlmPart.Image in generate. False = text-only.
Functions
Delegates to the SAME backend generateStream would pick — the first available one — so an app asking "what can I actually do right now" gets the answer for the backend it would really run against, not a generic composite-wide guess. No backend available reads the same as generate's empty-chain case: AiFailure.NotSupportedOnPlatform.
Multimodal entry point. Default maps a single LlmPart.Text onto generate (String) so every existing text-only actual (MediaPipe, Foundation Models, jvm) keeps working with zero changes. Backends that accept images (ML Kit GenAI) override this directly.
Delegates to the first available (and OnDeviceLlm.supportsImage-accepting) backend's OWN OnDeviceLlm.generateStream — not the single-emission default this class would otherwise inherit — so a backend with real token streaming (ML Kit GenAI, MediaPipe) actually streams through the one OnDeviceLlm every app gets from onDeviceLlmModule(), and cancelling the collecting coroutine reaches that backend while it's mid-generation.
Cheap, synchronous floor — true only when the platform could run inference. NOT the authoritative runtime check (model residency is async); generate still guards internally.