CloudOnDeviceLlm
OnDeviceLlm backed by an :llm-chat AiProvider chain — cloud as the on-device fallback tier. A caller appends one of these to CompositeOnDeviceLlm's backend list only when at least one API key exists, so desktop, web (wasmJs, which ships no real backend of its own — see OnDeviceLlm.wasmJs.kt) and any non-Nano/non-Foundation-Models Android or iOS device still gets a real answer through the same seam, instead of every app hand-maintaining its own always-false stub.
providers is expected to already be com.siddharth.kmp.llmchat.buildProviderChain's output (or a hand-built list of raw providers, in which case guarding them is the caller's job — this class assumes whatever list it's given is already guard-wrapped, same as CompositeOnDeviceLlm assumes of its own backends). isAvailable is providers.isNotEmpty(): a cheap synchronous floor per OnDeviceLlm's contract ("NOT the authoritative runtime check") — a real reachability check needs a network call, so that happens lazily inside generate/generateStream.
Properties
True when this backend accepts an LlmPart.Image in generate. False = text-only.
Functions
Honest self-report for a capability-aware caller: whether this backend genuinely streams tokens and accepts images, which GenerationConfig fields it reads, and — when it can't run — the real AiFailure reason instead of a bare false. Default answer is conservative (no streaming, no honored config) and derives AiCapabilities.unavailableReason from isAvailable alone; a backend with a more specific gate (model residency, OS feature status) overrides this to report why.
Detection-ordered, same as CompositeOnDeviceLlm.tryBackends: tries each provider in providers order, returns the first success, and the last failure when every provider declines or errors. AiFailure.NoKey when providers is empty — this device has no cloud key configured at all, distinct from a configured-but-failing provider.
Multimodal entry point. Default maps a single LlmPart.Text onto generate (String) so every existing text-only actual (MediaPipe, Foundation Models, jvm) keeps working with zero changes. Backends that accept images (ML Kit GenAI) override this directly.
Streams the first available provider's real SSE token stream (all three HTTP cloud providers genuinely stream — see :llm-chat's httpCloudCapabilities).
Cheap, synchronous floor — true only when the platform could run inference. NOT the authoritative runtime check (model residency is async); generate still guards internally.