🤗 Optimum Intel: Accelerate inference with Intel optimization tools
-
Updated
Aug 21, 2026 - Jupyter Notebook
🤗 Optimum Intel: Accelerate inference with Intel optimization tools
Inference engine for Intel devices. Serve LLMs, VLMs, Whisper, Kokoro-TTS, Embedding and Rerank models over OpenAI endpoints.
Just practising deploying MLM locally on my laptop
ov-cli — A local inference tool for LLMs based on OpenVINO. It supports model conversion (FP32/FP16/INT8/INT4), quantization, interactive chat (streaming output), and translation. It automatically recognizes both GenAI and Optimum formats, ready to use out of the box.
An experimental OpenAI-compatible API server optimized for LLM inference on Intel Integrated GPUs (iGPU)
A lightweight local model server that delivers an ollama‑style developer experience, built natively on Intel’s OpenVINO runtime for efficient CPU and GPU inference.
Model-agnostic GTK control center and OpenAI-compatible OpenVINO server for local AI models
Self-contained OpenVINO GenAI inference server with an OpenAI-compatible API.
OpenVINO environment for Intel on Linux.
Benchmarks and tuned OpenVINO configurations for Ornith-1.0-9B on Intel Arc 140T
Add a description, image, and links to the openvino-genai topic page so that developers can more easily learn about it.
To associate your repository with the openvino-genai topic, visit your repo's landing page and select "manage topics."