New Model Available: Ling-3.0-flash-Sante InclusionAI released Ling-3.0-Flash-Sante, a Mixture-of-Experts model built on Ling-3.0-Flash with 124 billion total parameters that activates approximately 5.1 billion parameters per token. The model targets health and medicine, with stated strengths in medical knowledge reasoning, clinical safety, evidence-based retrieval, and long-horizon medical tasks, while retaining general capabilities in reasoning, coding, and agentic tasks. Ling-3.0-Flash-Sante is a Mixture-of-Experts MoE model with enhanced capabilities across health and medicine. Built on Ling-3.0-Flash, it has 124 billion total parameters and activates approximately 5.1 billion parameters per token. The model excels in medical knowledge reasoning, clinical safety, evidence-based retrieval, and long-horizon medical tasks, while retaining strong general capabilities in reasoning, coding, and agentic tasks. Back to Models https://zenmux.ai/models Providers Route requests across multiple providers. Copy a provider slug to set your preference. ~~$0.06~~ $0 / M tokens ~~$0.18~~ $0 / M tokens Read: ~~0.012~~ 0 / M tokens Write: - / M tokens262.14K-- Uptime 24hours Direct request success rate on AI Gateway and per-provider. Throughput 24hours P50 throughput on live AI Gateway traffic, in tokens per second TPS . Latency 24hours P50 time to first token TTFT on live AI Gateway traffic, in milliseconds. Activity Token volume and request traffic to this model over time. Benchmarks Scores on standardized evaluations. Higher percentages are better — and rank percentile shows Metrics sourced from Artificial Analysis https://artificialanalysis.ai/ Apps Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for. View All https://zenmux.ai/analytics/apps Related Models More models from inclusionAI https://zenmux.ai/inclusionai