Skill
MLX VLM local vision inference
Native Apple-Silicon VLM inference via MLX-VLM: zero-server, vision feature caching for 11x batch speedup, CLI + Python API + FastAPI server + LoRA fine-tuning
Primitives inside (5)
deepseek-ocr-normalized-coordsreferenceDeepSeek-OCR(-2) grounding output uses coordinates normalized to 0-1000 (not pixels) in <ref>text</ref><box>[[x1,y1,x2,y2]]</box> — scale by image_dimension/1000 before drawing or cropping; the <|grounding|> prompt prefix triggers document-to-markdown mode.
When: Post-processing DeepSeek-OCR output into pixel-space boxes, crops, or markdown documents.
mlx-vlm-generate-returns-dataclassgotcha-fixmlx_vlm.generate() returns a GenerationResult dataclass, not a str — read .text for the answer (with free stats in .generation_tps/.peak_memory/.cached_tokens); str(result) yields the full dataclass repr.
When: Consuming generate() output programmatically — writing answers to files, feeding downstream prompts, parsing.
mlx-vlm-lora-flag-namesgotcha-fixmlx_vlm.lora takes --model-path and --dataset — NOT --model / --data as older docs show; wrong flags fail at argparse (verified against v0.6.4 --help).
When: Launching VLM LoRA fine-tuning with python -m mlx_vlm.lora.
mlx-vlm-memory-fit-calibrationcalibrationUnified-memory sizing for MLX-VLM on Apple Silicon: 8 GB fits 2B-4bit class (Qwen2-VL-2B ~4-5 GB), 16 GB fits 7B-4bit (~6 GB) and DeepSeek-OCR-2-bf16 (~7 GB), 32 GB+ needed for 27B/72B-4bit class (~40+ GB at the top).
When: Choosing a VLM/OCR model for a specific Mac before downloading weights.
mlx-vlm-vision-cache-opt-ingotcha-fixWhen asking several questions about the same image with MLX-VLM, keep them in ONE process/session AND make the vision-feature cache actually engage: in the Python API the cache is OPT-IN — plain generate() never caches — so create ONE VisionFeatureCache for the whole batch and pass vision_cache=cache on every call (LRU, keyed by image path or PIL content hash); only CLI --chat mode and the server's --vision-cache-size create it automatically; with the cache engaged repeat queries of the same image run ~11x faster than re-encoding per question.
When: Multi-question interrogation of a single image (QA gates, structured extraction via question lists, per-image QA over generated assets) on Apple Silicon with MLX-VLM.
Get the whole skill
All 5 primitives of this skill as one package, with the order to apply them.
Buy only the primitives you need
Each primitive is 1 credit (≈ €0.10). Pick them from the list above — the button is next to each one.
Upgrade your own skill
Paste your skill; we pick the 5 primitives from the shelf that fit it best, as one bundle for 5 credits (≈ €0.50).
Upgrade my skillNeighbour skills
Skills whose primitives are closest to this one (bge-m3 similarity):