Llama.cpp vs vLLM vs SGLang
Glad Labs decided not to switch its self-hosted inference stack from Ollama to vLLM after reviewing its own call logs, which showed only one to three concurrent calls at most against roughly 50 calls a day on an ASUS ROG…
Large language model (LLM) news — GPT-4, Claude, Gemini, Llama, Mistral and the latest research on training, fine-tuning, RLHF, and deployment of LLMs.
Glad Labs decided not to switch its self-hosted inference stack from Ollama to vLLM after reviewing its own call logs, which showed only one to three concurrent calls at most against roughly 50 calls a day on an ASUS ROG…
Few-shot prompting with three to five high-quality examples can fix classification errors that zero-shot prompts produce in LLM applications, according to a developer account published on PromptCube. The author reported …
LangChain's LangSmith observability platform can record every step of non-deterministic AI agent workflows by organizing data into OpenTelemetry-like runs and trace trees, according to a technical walkthrough published b…
An analysis of the HTTP Archive's 1 August 2026 mobile crawl, covering 17.2 million pages, found that about 32% of pages serve incomplete content to non-rendering AI crawlers, with roughly 9% delivering under half their …
Large language models have no persistent personal memory of past interactions, and any apparent recall is handled by the application built around the model rather than the model itself, according to a technical explainer…
A developer benchmarked Z.ai's GLM 5.3 and GLM 5.3 Flash against GLM 5.2 and DeepSeek V4.1 Flash, finding that GLM 5.3 no longer allows thinking to be disabled — requests with thinking disabled or reasoning_effort set to…
DeepSeek's Coder V2 and V3 models are significantly cheaper than Anthropic's Claude 3.5 Sonnet, often by a factor of 10x or more, but Claude 3.5 Sonnet remains more reliable for complex software engineering and large-con…
A participant in the Recurse Center summer programming retreat in Brooklyn documented work across four study groups, including Agentic Adventures, where members built a remote sandbox for a "vibecoding agent" run with da…
A developer published a step-by-step guide for making a first LLM API call in Python, walking through setting up a virtual environment, installing the Anthropic SDK and python-dotenv, and calling Claude via the messages.…
A developer demonstrated a setup for running the open-source Hermes Agent harness from Nous Research on free AI models by inserting OmniRoute, an open-source routing gateway, between the agent and model providers. The ar…
Frontier AI labs have a financial incentive to advocate for regulatory pacing that protects their commercial position, according to an essay published at pacingthefrontier.com. The essay cites Anthropic CEO Dario Amodei'…
A developer named Bhuvanesh selected Amazon Bedrock for an AWS rhyming game contest, producing an overview of the fully managed generative AI service. The writeup explains how Bedrock provides access to foundation models…
Anthropic CEO Dario Amodei called for the AI industry to "pace the frontier," slowing the rate of capabilities advancement so risk prevention can keep up, a proposal quickly embraced by Sam Altman, Elon Musk, and Demis H…
A feature request proposes adding keyboard shortcuts to the Assistant for switching between AI models, such as Cmd/Ctrl + 1 for ChatGPT 4o and Cmd/Ctrl + 2 for Claude 3.5 Sonnet, matching the effect of the model dropdown…
A forum poster at Level1Techs described being asked to roll out AI tools across a 30-employee company where some staff already hold personal ChatGPT subscriptions, and outlined a plan weighing Microsoft 365 Copilot at $7…
A developer prompted DeepSeek to design a band-pass filter with a 1 kHz lower cutoff, 10 kHz upper cutoff, and a gain of 50 using a 741 op-amp. The model produced a clean two-stage design and flagged the 741's 1 MHz gain…
Salesforce introduced Koa, its first CRM reasoning model for Agentforce, built on NVIDIA Nemotron open models and trained on 27 years of Salesforce CRM intelligence. Salesforce reports Koa is 11% more precise at calling …
A Hacker News user posting as tonyvince7 asked why large language models consistently pick 17 when prompted to choose a number between 1 and 30, saying the behavior appeared across multiple models. The Ask HN post, submi…
Gavel, a method that extracts skill-routing signals from a frozen LLM's hidden states using two trained linear maps, outperforms progressive disclosure and retrieve-and-rerank pipelines adding 1.2B to 16B external parame…
Researchers submitted a paper to arXiv on 5 Feb 2026 introducing GRP-Obliteration (GRP-Oblit), a method that uses Group Relative Policy Optimization (GRPO) to remove safety constraints from aligned models using a single …