Package-level declarations
Types
Preferred compute unit for on-device inference. A hint, not a guarantee: a backend maps it to the nearest unit its engine offers and degrades silently — MediaPipe (tasks-genai) has no NPU backend so NPU falls back to the engine default, and Gemini Nano always runs on the NPU/AICore regardless.
OnDeviceLlm backed by an :llm-chat AiProvider chain — cloud as the on-device fallback tier. A caller appends one of these to CompositeOnDeviceLlm's backend list only when at least one API key exists, so desktop, web (wasmJs, which ships no real backend of its own — see OnDeviceLlm.wasmJs.kt) and any non-Nano/non-Foundation-Models Android or iOS device still gets a real answer through the same seam, instead of every app hand-maintaining its own always-false stub.
Detection-ordered OnDeviceLlm: probes an ordered list of backends and uses the first one that both reports available AND actually produces output. This is how a device escalates ML Kit Gemini Nano (AICore-only) → MediaPipe Gemma (broad device coverage, downloaded on demand) → CloudOnDeviceLlm (network fallback, appended by the caller only when a key exists) → (nothing → the caller's own heuristic fallback tier).
Live progress of a model download. receivedBytes/totalBytes drive a progress bar; bytesPerSec and etaMs drive a "12 MB/s · 2 min left" label. totalBytes and etaMs are -1 when unknown (server sent no Content-Length).
Real actual: Apple Foundation Models, reached via a Swift class conforming to NativeLlm and registered into FoundationModelsBridge at app startup — see ai/ios-bridge/README.md. Kotlin/Native has no platform.FoundationModels.* cinterop binding (the framework's Swift-macro-driven @Generable/streamResponse surface has no ObjC-compatible shape), so this follows the same bridge mechanism Doori already shipped for FoundationModelsAnalyzer/ FoundationModelsLlmGateway: export a plain Kotlin interface to Swift as an ObjC protocol, implement it in Swift, inject the implementation through a top-level singleton at startup.
Optional load-time tuning for an on-device LLM. Every field is nullable — null means "leave the backend's own default alone". Backends apply what their engine exposes and ignore the rest:
Generic delegate-or-degrade injection seam, same shape as Doori's InjectableDocumentAiAnalyzer/ InjectableTextGenerator — kept in commonMain (not iosMain) so this logic is unit-testable; the iosMain holder (FoundationModelsBridge) has nothing left to test once it just forwards here.
Buckets free text into one of a fixed set of categories by keyword hits — no model call, no network, no on-device inference. The "bucket this free text" case doesn't always need a full StructuredOutput round trip; when the categories are known in advance (intent routing, ticket triage), this is a shared, tested classifier instead of every caller hand-rolling its own .contains() chain that breaks on a synonym.
One piece of multimodal input to OnDeviceLlm.generate. ByteArray (not a platform bitmap type) keeps this commonMain-safe — each platform actual decodes the bytes itself.
Manages the on-demand MediaPipe Gemma model file in app-private storage. The model binary is NEVER committed to the repo — download is user-triggered (surfaced on the settings screen via ModelManager) and lands in filesDir/models/, resuming from a .tmp if a prior attempt was cut off.
Optional Gemini Nano backend; supplied by :ai-mlkit on Play-enabled builds.
Lifecycle of a downloadable on-device model (e.g. MediaPipe Gemma). PARTIALLY_DOWNLOADED marks a .tmp left behind by an interrupted download — the next ModelManager.download resumes from it.
A downloadable on-device model, surfaced to the settings screen.
The settings-screen-facing control surface for optional downloadable on-device models. Kept tiny on purpose (list / observe / download / delete). Backends that need no download (ML Kit Gemini Nano is managed by AICore; Foundation Models by the OS) don't appear here — only models the app fetches itself (MediaPipe Gemma) do.
Config-driven description of a downloadable on-device model — the manifest a ModelManager reads instead of hard-coding one model. Keeps model choice out of code: an app ships (or fetches) a list of these and the manager downloads/manages each by id.
The seam a platform's native (non-Kotlin/Native-importable) LLM API implements and gets injected through, so a Kotlin actual (here, FoundationModelsOnDeviceLlm) reaches it without a cinterop binding. Hoisted from Doori's hand-copied InjectableDocumentAiAnalyzer/InjectableTextGenerator shape (core:ai/feature:agent there) so the pattern exists once in the toolkit instead of being re-copied into every consuming app.
Cancels an in-flight NativeLlm.generateStream call. Safe to call more than once.
Delivered from the native side as NativeLlm.generateStream progresses. Exactly one of onComplete/onError fires, always after the last onPartial.
Default for platforms/targets with no downloadable models (JVM/desktop, iOS today).
A single-shot on-device text LLM tier. Kept deliberately tiny — one availability gate and one text-in/text-out call — so each platform's actual (ML Kit GenAI on Android, Foundation Models on iOS, unavailable elsewhere) is a thin wrapper, and DefaultJobIntelligence never has to know which backend ran.
Turns "parse a typed T out of a model's free-text reply" from a hand-rolled regex scrape (the kind that breaks the first time a model adds a newline or an explanatory sentence) into one shared, tested path: a schema hint embedded in the prompt, a tolerant JSON parse of the reply, one repair retry that shows the model its own unparseable output, and a typed AiFailure otherwise.
The common fallback tier: no on-device model. Desktop/JVM/wasm and any pre-AI device land here.
Marks an OnDeviceLlm that compiles and satisfies the seam but has no real implementation behind it yet — it reports unavailable unconditionally rather than attempting partial/incorrect work. Distinct from UnavailableOnDeviceLlm (a deliberate, permanent "no model on this platform/target" answer): a class carrying this annotation WOULD work once reason (its missing native bridge) is built — see the annotated class's own upgrade-path doc comment.
Properties
Functions
Pure progress/speed/ETA calc for a resumable download — no I/O, so it is unit-testable. received is the total bytes on disk (including any resumed startOffset); elapsedMs is time since THIS session started, so speed reflects the live transfer, not the resumed head start.
Android on-device LLM tier, detection-ordered (ai-engineering.md §7): ML Kit Gemini Nano (AICore devices) → MediaPipe Gemma (broad coverage, downloaded on demand) → (falls through to the heuristic tier upstream). ModelManager is bound for the settings screen. Gemini Nano participates only when the consumer installs :ai-mlkit and mlKitLlmModule().
Per-platform Koin bindings for the on-device LLM tier. commonMain's aiModule includes this; the actual decides which OnDeviceLlm gets bound (ML Kit / Foundation Models / unavailable).
iOS on-device LLM tier, detection-ordered: Apple Foundation Models → MediaPipe Gemma → (falls through to the heuristic tier upstream).
Desktop/JVM has no on-device model — the heuristic tier always answers.
Web/wasmJs has no on-device model — same floor as OnDeviceLlm.jvm.kt. A caller wanting a real answer in the browser wires its own CloudOnDeviceLlm (an :llm-chat com.siddharth.kmp.llmchat.AiProvider chain reaches every cloud vendor's HTTP API fine from wasmJs) instead of this module shipping one itself — this module owns no API keys.
Convenience for the common case: build a StructuredOutput from a reified @Serializable type, no explicit .serializer() call at the use site.