Skip to main content
← Back to Vision Models LFM2.5-VL-3B is Liquid AI’s most capable vision-language model, with strong grounding, screen and document understanding, and function calling. It builds on the LFM2.5-2.6B backbone with a SigLIP2 NaFlex image encoder, and answers directly for low-latency inference on-device and in the cloud.

Specifications

Grounding & Detection

Object detection and localization from natural-language queries

Screen & Document Understanding

Digital screens, full-page OCR, and layout-aware parsing

Function Calling

Tool use from text-only and vision-text inputs

Quick Start