grok-4.3 edges gpt-5.4-mini on execution
Grok 4.3 outperformed GPT 5.4 Mini in a head-to-head execution benchmark, scoring 38.3 to 36.2 by demonstrating greater reliability on formatting, tone control, and frictionless output. In a key test …
Grok 4.3 outperformed GPT 5.4 Mini in a head-to-head execution benchmark, scoring 38.3 to 36.2 by demonstrating greater reliability on formatting, tone control, and frictionless output. In a key test …
Coram AI, founded by Ashesh Jain and Peter Ondruska, raised $35 million in Series B funding to expand its AI platform that makes existing security cameras searchable. The round, co-led by Ansa Capital…
Wharton professor Ethan Mollick reported on June 9 that Anthropic's Claude Fable 5 model executed multi-page specifications for up to a dozen hours, outperforming public models he had used in sustaine…
Anthropic CEO Dario Amodei called on Washington to regulate the most powerful AI models under an FAA-style safety regime that could delay, block, or reverse releases failing public safety tests. The p…
Seedance 2 Image to Video scored 17.0 against AnimateDiff's 6.7 in a direct comparison of prompt fidelity, decisively winning on the ability to translate complex text into specific visual sequences. T…
Guan Wang, CEO of Sapient Intelligence, claims his researchers trained a 1 billion parameter language model from scratch for approximately $1,500, according to a June 10 VentureBeat report. The model,…
Igor Babuschkin, a co-founder of Elon Musk's xAI, has launched a new company focused on personalized artificial intelligence, Bloomberg reported on June 10. The venture shifts Babuschkin away from the…
Imagineart 2.0 Preview outperformed AuraFlow in a direct comparison by delivering superior results in prompt fidelity, scene logic, and typography, the three key metrics of image model usefulness. Whi…
Ari Jacoby's Concentrate AI launched from stealth Wednesday with over $5 million in funding, entering the competitive AI routing market dominated by OpenRouter. The company aims to serve as a control …
Grok-4.3 scored 33.8 against GPT-5.4 Nano's 33.4 in a head-to-head evaluation, with the split revealing distinct strengths. GPT-5.4 Nano outperformed in writing tasks, including Python log redaction a…
Yoni Ramon and Guy Arazi launched Pi from stealth Wednesday with $35 million in funding, valuing the AI security startup at $100 million. Pi uses an AI agent that analyzes a customer's code, policies,…
Anthropic Chief Product Officer Mike Krieger used a thread on X to frame the company's June 9 launch of Claude Fable 5 as a product test, asserting the model can handle longer delegated tasks while th…
Instawork CEO Sumir Meghani launched Instacore, a wearable camera system that allows hourly workers to record commercial tasks for robotics companies to train AI models, Business Insider reported June…
Accessibility communications service Nagish has rebranded as Rylo and raised $85 million in growth funding, the company announced on June 9. The round, led by General Catalyst and Canaan alongside exi…
A purported leaked screenshot of DeepSeek's graphical user interface shows an agent-first workspace layout that closely mirrors OpenAI's Codex interface, though the image remains unverified. The leak,…
Clive Chan, OpenAI's second hardware hire who worked on the company's custom chip program, has left the company and joined Anthropic this week, according to a LinkedIn post. Chan praised his former te…
CJ Zafir teased Mac-1, a 6.6 billion parameter AI model designed to run locally on Macs and operate native macOS tools, in a thread on X on Saturday. Zafir claims the model requires 7GB of RAM, runs b…
GPT Image 2 API defeated AuraFlow in a direct comparison, winning every task with a final score of 27.5 to 18.6. The API demonstrated superior precision by accurately rendering specific prompt details…
Lockheed Martin’s AI Center hosted its inaugural AI Fight Club event, testing combat AI agents in a synthetic environment designed to simulate high-pressure military decision-making for joint all-doma…
Patrick Jiang (@patpcj) released Harness-1, a 20 billion parameter search agent that externalizes search state from the model into a structured harness, as detailed in a 13-post thread on X and an acc…