generateStream
Delegates to the first available (and OnDeviceLlm.supportsImage-accepting) backend's OWN OnDeviceLlm.generateStream — not the single-emission default this class would otherwise inherit — so a backend with real token streaming (ML Kit GenAI, MediaPipe) actually streams through the one OnDeviceLlm every app gets from onDeviceLlmModule(), and cancelling the collecting coroutine reaches that backend while it's mid-generation.
// ponytail: no cross-backend fallback once a stream is chosen — unlike generate, which can // retry the next backend after a clean failure, a stream that has already shown the user a // few tokens can't un-show them, so escalating mid-stream would look like a glitch, not a // fallback. Only the pre-flight choice of WHICH backend to stream from escalates. Add // real fallback (buffer until the first emission proves the backend live) if an empty/failed // stream turns out to be common enough in practice to matter.