--- library_name: llima license: other license_name: lfm1.0 license_link: LICENSE tags: - vision - image-text-to-text - generative_ai - embedded - sima - liquidai - lfm2 pipeline_tag: image-text-to-text base_model: LiquidAI/LFM2-VL-3B --- # LFM2-VL-3B-Autoround-a16w4: Optimized for SiMa.ai Modalix ## Overview This repository contains the **LFM2-VL-3B-Autoround-a16w4** model, optimized and compiled for the **SiMa.ai Modalix** platform. - **Model Architecture:** LFM2-VL (3B parameters) - **Quantization:** Hybrid - **Pre-quantization:** The source checkpoint was pre-quantized with AutoRound before compilation. - **Prompt Processing:** A16W8 (16-bit activations, 8-bit weights) - **Token Generation:** A16W4 (16-bit activations, 4-bit weights) - **Maximum context length:** 8192 - **Input Resolution:** 512x512 (Fixed) - **Source Model:** [LiquidAI/LFM2-VL-3B](https://huggingface.co/LiquidAI/LFM2-VL-3B) ## Performance ### Multimodal inference The following measurements use one 512x512 image at the model's compiled input resolution and a 20-token text question, so no image resizing is required. Measured end-to-end on Modalix using five fresh requests after warm-up. Values are arithmetic means. TTFT includes image decoding and preprocessing, vision encoding and projection, language-model prefill, and the first generated token. Generation rate is measured after the first token. | Image input | Text tokens | Mean TTFT (seconds) | Mean generation rate (tokens/second) | |---:|---:|---:|---:| | 512x512 | 20 | 0.41 | 38.13 | ### Text-only generation The following measurements exercise the language-model component without an image. Vision encoding and projection latency are not included, and input tokens refer only to text tokens. Measured with MoLE using batch size 1, five samples per input length, and up to 128 generated tokens. Values are arithmetic means. TTFT includes language-model prefill and the first generated token; generation rate is measured after the first token. | Input tokens | Mean TTFT (seconds) | Mean generation rate (tokens/second) | |---:|---:|---:| | 128 | 0.08 | 38.25 | | 256 | 0.15 | 38.23 | | 512 | 0.30 | 37.89 | | 1024 | 0.61 | 37.30 | | 2048 | 1.28 | 35.28 | | 3072 | 2.16 | 34.37 | | 4096 | 3.13 | 33.13 | | 5120 | 4.38 | 32.17 | | 6144 | 5.76 | 31.82 | | 7168 | 7.57 | 30.67 | ## Prerequisites To run this model, you need: 1. **SiMa.ai Modalix Device** 2. **SiMa.ai CLI**: Installed on your Modalix device. 3. **SiMa.ai Neat Runtime**: Install or update the Neat Library on Modalix. The LLiMa runtime is installed as part of the Neat runtime. 4. **Hugging Face CLI**: Optional, for downloading the model on a host before copying it to Modalix. ## Installation & Deployment Follow these steps to deploy the model to your Modalix device. ### 1. Install or Update Neat Runtime > **Note:** This is a **one-time setup**. If the Neat Library is already installed on your Modalix device, you can skip this step and continue with model download. Follow the [SiMa.ai Neat getting started guide](https://developer.sima.ai/software/getting-started/) to install or update the Neat Library on your Modalix device. The `llima` CLI is available on Modalix after the Neat runtime is installed. It manages precompiled GenAI models under `/media/nvme/llima/models` by default. Set `LLIMA_MODELS_PATH` to use a different model directory. ### 2. Download the Model Download the compiled model assets from this repository directly to your device. ```bash # Download the model to a local directory llima pull LFM2-VL-3B-Autoround-a16w4 ``` Alternatively, you can download the compiled model to a Host and copy it to the Modalix device: ```bash hf download simaai/LFM2-VL-3B-Autoround-a16w4 --local-dir LFM2-VL-3B-Autoround-a16w4 scp -r LFM2-VL-3B-Autoround-a16w4 sima@:/media/nvme/llima/models/ ``` *Replace \ with the IP address of your Modalix device.* **Expected Directory Structure:** ```text /media/nvme/llima/ └── models/ └── LFM2-VL-3B-a16w4/ # The compiled model ``` ## Usage ### Validate with LLiMa CLI Run the model directly on Modalix: ```bash llima run LFM2-VL-3B-Autoround-a16w4 ``` For all runtime options, run: ```bash llima run -h ``` ### GenAI Demo Application The GenAI demo application is separate from LLiMa installation. Use the [GenAI Multimodal Assistant](https://developer.sima.ai/examples/app/genai%2Fmultimodal-assistant) page to install and run the demo app. Once installed, the demo app can use precompiled models such as this one. ### API Usage To serve this model with OpenAI- or Ollama-compatible APIs and send requests to it, use the GenAI server workflow in [Serve GenAI Models](https://developer.sima.ai/software/tutorials/serve-genai-models). For direct VLM calls without setting up a server, see [Run a VLM](https://developer.sima.ai/software/tutorials/run-a-vlm). ## Limitations - **Quantization**: This model is quantized (A16W4/A16W8) for optimal performance on embedded devices. While this maintains high accuracy, minor deviations from the full-precision model may occur. - **Fixed Resolution**: While the standard LFM2-VL architecture supports dynamic input resolutions, this version has been specifically optimized and fixed to **512x512** resolution at compile time to achieve maximum throughput and efficiency on the SiMa.ai MLA. ## Troubleshooting - **`sima-cli` not found**: Ensure that `sima-cli` is installed on your Modalix device. - **`llima` not found**: Install or update the Neat Library. See [Getting Started](https://developer.sima.ai/software/getting-started/). - **Model can't be run**: Verify the model directory is exactly inside `/media/nvme/llima/models/` and not nested (e.g., `/media/nvme/llima/models/LFM2-VL-3B-Autoround-a16w4/LFM2-VL-3B-Autoround-a16w4`). - **Permission Denied**: Ensure you have read/write permissions for the `/media/nvme` directory. ## Resources - [GenAI with LLiMa](https://developer.sima.ai/software/genai-llima/) - [Serve GenAI Models](https://developer.sima.ai/software/tutorials/serve-genai-models) - [Run a VLM](https://developer.sima.ai/software/tutorials/run-a-vlm) - [GenAI Multimodal Assistant](https://developer.sima.ai/examples/app/genai%2Fmultimodal-assistant)