cd/sources/hugging-face-blog· home› sources› Hugging Face Blog
cat /sources/hugging-face-blog.feed | wc -l → 1013

Hugging Face Blog

articles 1013 domain huggingface.co → page 51/51 feed RSS
00:00
2026-05-14
huggingface.co
machine-learning

Unlocking asynchronicity in continuous batching

Synchronous continuous batching in LLM inference causes inefficiency by forcing the CPU and GPU to work sequentially, leaving one idle while the other operates. This idle time can account for nearly a…

23:18
2026-05-11
huggingface.co
artificial-intelligence

Building Blocks for Foundation Model Training and Inference on AWS

Convergent infrastructure requirements for the foundation model lifecycle—including tightly coupled accelerator compute, high-bandwidth networking, and distributed storage—and highlights the growing r…

19:06
2026-05-06
huggingface.co
large-language-models

vLLM V0 to V1: Correctness Before Corrections in RL

Here is a 2-3 sentence factual summary of the article: The article describes the process of migrating an online reinforcement learning (RL) training system from the vLLM V0 engine to the V1 rewrite, …

00:00
2026-05-06
huggingface.co
artificial-intelligence

Adding Benchmaxxer Repellant to the Open ASR Leaderboard

The Open ASR Leaderboard has added private, high-quality English ASR datasets from Appen Inc. and DataoceanAI to prevent "benchmaxxing" and test-set contamination. While the average Word Error Rate (W…

15:01
2026-04-29
huggingface.co
large-language-models

Granite 4.1 LLMs: How They’re Built

The Granite 4.1 family consists of dense, decoder-only LLMs (3B, 8B, and 30B parameters) trained from scratch on approximately 15 trillion tokens through a five-phase pre-training pipeline that progre…

00:00
2026-04-29
huggingface.co
artificial-intelligence

DeepInfra on Hugging Face Inference Providers 🔥

DeepInfra has been added as a supported Inference Provider on the Hugging Face Hub, offering serverless AI inference with over 100 models and cost-effective per-token pricing. The integration initiall…

00:00
2026-04-27
huggingface.co
artificial-intelligence

How to build scalable web apps with OpenAI's Privacy Filter

Three scalable web applications—a Document Privacy Explorer, an Image Anonymizer, and a SmartRedact Paste tool—all built using OpenAI's Privacy Filter model and Gradio's Server infrastructure. The Pri…

00:00
2026-04-24
huggingface.co
large-language-models

DeepSeek-V4: a million-token context that agents can actually use

DeepSeek-V4 introduces a new architecture using hybrid attention mechanisms—Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA)—to drastically reduce the computational cost and me…

00:00
2026-04-23
huggingface.co
developer-tools

How to Use Transformers.js in a Chrome Extension

Technical guide for developers on integrating Transformers.js into a Chrome extension under Manifest V3 constraints. It details a three-part architecture using a background service worker to host AI m…

00:00
2026-04-21
huggingface.co
artificial-intelligence

AI and the Future of Cybersecurity: Why Openness Matters

The AI system Mythos, which combines a large language model with substantial compute power, scaffolding, and autonomy, can rapidly find and patch software vulnerabilities, highlighting that the system…

← prev page 51 / 51