Skip to main content
← Back to Decision Models d1-omni-600M is a 587M parameter decision model built on LFM2.5-Encoder-350M with a SigLIP2-based vision encoder and audio encoder jointly fine-tuned for single-pass decisions. It evaluates text with images or text with audio and returns typed answers with zero output tokens.

Specifications

Multimodal

Text + images or text + audio (up to 30 s)

Zero Output Tokens

Answers read from the distribution, no generation

Edge-Sized

587M parameters for on-device deployment

Quick Start

Install:
Text input:
Image input:
Audio input (16 kHz mono, 0.5–30 s):