cd/entity/HuggingFace· home› entities› HuggingFace
grep -l @huggingface /news/*.json | wc -l → 252

HuggingFace

mentions 252 type Organization page 10/13 feed RSS

// recent coverage 252 mentions

11:02
2026-06-30
ronreiter.com
large-language-models

Coding with DeepSeek 4 on a 128GB MacBook Pro

DeepSeek V4 Flash, a 284-billion-parameter Mixture-of-Experts model, now runs locally on a 128GB MacBook Pro via antirez's experimental llama.cpp fork, achieving ~21 tokens/sec generation on the Metal…

05:32
2026-06-29
squish.run
large-language-models

Squish – The fastest way to run local LLMs on Apple Silicon

Squish, a new local AI agent runtime for Apple Silicon, claims to load models 54× faster than standard paths and serve them faster than Ollama, with full offline capability and no cloud dependencies. …

00:00
2026-06-29
jasonrobert.dev
large-language-models

News Summary for June 29, 2026

DeepReinforce released Ornith-1.0, an open-source family of agentic coding models that use self-scaffolding reinforcement learning to outperform Claude Opus 4.7 on SWE-Bench Verified and Terminal-Benc…

07:08
2026-06-28
byteiota.com
large-language-models

GLM-5.2 Open Source: 750B Params, MIT License, 1M Context

Z.ai open-sourced GLM-5.2 on June 17 under an MIT license, a 744B-parameter sparse MoE model with a 1M-token context window that outperforms GPT-5.5 on multiple coding benchmarks while costing about o…

02:22
2026-06-26
lesswrong.com
ai-safety

Research note on negated reward hacking

Researchers at BlueDot's Technical AI Safety Project Sprint found that fine-tuning language models on negated documents can still teach them reward-hacking knowledge, leading to emergent misalignment …

17:36
2026-06-25
devclubhouse.com
large-language-models

Quantize and Run Llama 3.2 on Apple Silicon with llama.cpp

Mariana Souza published a tutorial on quantizing and running Meta's Llama 3.2 3B model on Apple Silicon using llama.cpp with Metal GPU acceleration, achieving local inference with Q4_K_M quantization.…

12:09
2026-06-25
byteiota.com
large-language-models

GLM-5.2 Beats GPT-5.5 at Coding for One-Sixth the Price

Z.AI's open-weight GLM-5.2 model outperforms GPT-5.5 on the SWE-bench Pro coding benchmark, scoring 62.1 versus 58.6, while costing $1.40 per million input tokens compared to GPT-5.5's $8.00. Released…

11:08
2026-06-25
flama.dev
large-language-models

LLM APIs with built-in chatbot in 1 line of code

Flama 2.0 introduces a CLI tool that allows users to download, package, and serve large language models from HuggingFace with a single command, including a built-in chat interface and production-ready…

13:01
2026-06-24
gist.github.com
large-language-models

NVIDIA GenAI LLM Certification Lab

NVIDIA has released a GenAI LLM Certification Lab that guides developers through building a production-ready fine-tuning and optimization pipeline. The lab covers data preparation, LoRA fine-tuning wi…

03:11
2026-06-24
byteiota.com
large-language-models

Baidu Unlimited-OCR: One-Shot PDF Parsing Is Here

Baidu released Unlimited-OCR, a new model that uses Reference Sliding Window Attention to parse up to 40 PDF pages in a single inference pass, eliminating the linear memory growth of traditional LLM-b…

15:34
2026-06-23
discuss.huggingface.co
machine-learning

Native binary embeddings experiment: curious about your thoughts

A developer tested native binary embedding training against post-hoc binarization using a small BERT-mini model and found that native training with a binary loss produced better retrieval results on S…

← prev page 10 / 13 next →
// co-occurs with top 8 entities
// topics top 6 topics