IBM: Granite 4.0 Micro
Long contextGranite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models.
OpenRouter-style model index with deep metadata: pricing, context windows, architecture, and parameters.
Granite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models.
Mercury is the first diffusion large language model (dLLM).
Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM).
Mercury Coder is the first diffusion large language model (dLLM).
Inflection 3 Pi powers Inflection's Pi chatbot, including backstory, emotional intelligence, productivity, and safety.
Inflection 3 Productivity is optimized for following instructions.
KAT-Coder-Pro V1 is KwaiKAT's most advanced agentic coding model in the KAT-Coder series.
KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integration.
LFM2 is a new generation of hybrid models developed by Liquid AI, specifically designed for edge AI and on-device deployment.
LFM2-24B-A2B is the largest model in the LFM2 family of hybrid architectures designed for efficient on-device deployment.
LFM2-8B-A1B is an efficient on-device Mixture-of-Experts (MoE) model from Liquid AI’s LFM2 family, built for fast, high-quality inference on edge hardware.
LFM2.5-1.2B-Instruct is a compact, high-performance instruction-tuned model built for fast on-device AI.
LFM2.5-1.2B-Thinking is a lightweight reasoning-focused model optimized for agentic tasks, data extraction, and RAG—while still running comfortably on edge devices.
An attempt to recreate Claude-style verbosity, but don't expect the same level of coherence or memory.
LongCat-Flash-Chat is a large-scale Mixture-of-Experts (MoE) model with 560B total parameters, of which 18.6B–31.3B (≈27B on average) are dynamically activated per input.