cd/entity/DeepSeek V4 Pro· home entities DeepSeek V4 Pro
grep -l @deepseek v4 pro /news/*.json | wc -l → 118

DeepSeek V4 Pro

mentions 118 type Person page 2/6 feed RSS

// recent coverage 118 mentions

18:13
2026-08-04
twitter.com
artificial-intelligence

Artificial Analysis Endpoint Accuracy Index

Artificial Analysis launched its Endpoint Accuracy Index, measuring how much of an open weights model's accuracy each serverless API endpoint preserves, with initial coverage of GLM-5.2, gpt-oss-120b,…

15:17
2026-08-04
dev.to
artificial-intelligence

DeepSeek V4 Flash API Cost: Thinking Mode Corrupts Strict JSON

DeepSeek's V4 Flash 0731 build corrupts integer fields in strict JSON schema outputs when thinking mode is enabled, failing 8 of 13 runs across two request paths, while disabling thinking fixes all ru…

13:09
2026-08-04
byteiota.com
artificial-intelligence

Thinking Machines Inkling: 975B Open-Weight Model for Fine-Tuning

Thinking Machines Lab released Inkling on July 15, a 975-billion-parameter Mixture-of-Experts model with 41B active parameters, open-weight under Apache 2.0, from Mira Murati's $12B startup. The model…

04:01
2026-08-04
trae1oung.github.io
artificial-intelligence

SWE-Touch: Benchmarking Coding Agents When Users Touch the Code

Researchers introduced SWE-Touch, a benchmark evaluating coding agents' ability to repair software after users directly edit code, finding that most models' performance drops significantly when user e…

14:31
2026-07-30
pub.towardsai.net
artificial-intelligence

I Tried to Catch 5 AIs Favoring Themselves. Only Some Did.

A test of five AI models — GPT-5.6 Sol, DeepSeek V4 Pro, Claude Fable 5, Grok-4.5, and Gemini 3.1 Pro — found that only two showed clear self-preference bias when scoring their own unlabeled essays al…

22:08
2026-07-29
byteiota.com
artificial-intelligence

DeepSeek V4 Pro: 80.6% SWE-Bench at $0.87/M Output

DeepSeek V4 Pro, a Mixture-of-Experts model with 1.6 trillion total parameters, scores 80.6% on SWE-bench Verified, tying Gemini 3.1 Pro and achieving the highest score for any open-weight model. Pric…

05:35
2026-07-25
dev.to
ai-agents

Building Verification Loops in Claude Code with Skills

Anthropic published a guide on building verification loops in Claude Code with skills, enabling AI agents to automatically check and fix their own work using tests, linters, and custom checks. The app…

05:35
2026-07-25
dev.to
large-language-models

SCOTOMA: Gemma 4 31B Abliteration Review

Developer Nokka reviewed SCOTOMA, an abliterated version of Google's Gemma 4 31B Instruct model created by ReadyArt. SCOTOMA uses a novel J-space abliteration technique with Jacobian-lens projection t…

13:46
2026-07-24
gist.github.com
developer-tools

Myt AI Console - OpenCode configuration

A developer shared a configuration guide for integrating Myt AI Console's custom AI models into OpenCode. The setup involves appending an opencode-config.json file to the global OpenCode profile and a…

04:10
2026-07-23
byteiota.com
artificial-intelligence

Kimi K3: What 2.8 Trillion Parameters Actually Buys You

Moonshot AI released Kimi K3 on July 16, a 2.8-trillion-parameter Mixture-of-Experts model that activates only ~50 billion parameters per token, making it the largest open-weight model ever built. K3 …

00:14
2026-07-23
github.com
ai-agents

PenguinHarness, Open, Efficient, Self-Improving Harness

PenguinHarness, an open-source agent framework from Prism Shadow, claims to build agents at 100× the speed of LangChain with a zero-code CLI and Web UI connected to 1000+ models, achieving best accura…

17:17
2026-07-22
dylancastillo.co
large-language-models

Are AI Labs Pelicanmaxxing?

Simon Willison's informal benchmark asking AI models to generate an SVG of a pelican riding a bicycle has become a widely discussed test for large language models. A new experiment tested 1,008 SVGs a…

13:08
2026-07-21
sourcefeed.dev
artificial-intelligence

LM Studio's Bionic Agent Is Also Its Cloud Pivot

LM Studio released Bionic on July 16, a new agent application for Mac and Windows that marks the company's first metered product with a checkout page, pivoting from its free local LLM desktop app to a…

← prev page 2 / 6 next →
// co-occurs with top 8 entities
// topics top 6 topics