Coding with DeepSeek 4 on a 128GB MacBook Pro
DeepSeek V4 Flash, a 284-billion-parameter Mixture-of-Experts model, now runs locally on a 128GB MacBook Pro via antirez's experimental llama.cpp fork, achieving ~21 tokens/sec generation on the Metal…
DeepSeek V4 Flash, a 284-billion-parameter Mixture-of-Experts model, now runs locally on a 128GB MacBook Pro via antirez's experimental llama.cpp fork, achieving ~21 tokens/sec generation on the Metal…
A new audit of over twenty African NLP corpus families reveals widespread license incompatibilities, including CC-BY-SA and CC-BY-NC datasets that cannot be legally combined and NoDerivs clauses that …
An open-source benchmark for prompt-injection detectors has been released, measuring both attack catch rate and false positives on real traffic in a threshold-agnostic way. The benchmark evaluates ten…
Squish, a new local AI agent runtime for Apple Silicon, claims to load models 54× faster than standard paths and serve them faster than Ollama, with full offline capability and no cloud dependencies. …
DeepReinforce released Ornith-1.0, an open-source family of agentic coding models that use self-scaffolding reinforcement learning to outperform Claude Opus 4.7 on SWE-Bench Verified and Terminal-Benc…
Moonshot AI released Kimi K2.7-Code, an open-source 1-trillion-parameter coding agent, on June 25, claiming 30% fewer thinking tokens and higher benchmark scores than its predecessor. The model uses a…
Z.ai open-sourced GLM-5.2 on June 17 under an MIT license, a 744B-parameter sparse MoE model with a 1M-token context window that outperforms GPT-5.5 on multiple coding benchmarks while costing about o…
Developer Rand01ph released GemmaTrans, an on-device translation app for macOS built with MLX-Swift, supporting Google Gemma 4 and Tencent Hy-MT2 models fully offline. The app provides a local HTTP AP…
Mixture-of-Experts (MoE) models like Qwen3-30B-A3B and DeepSeek-V3 separate total parameters (memory) from active parameters (compute), allowing a 30B-parameter model to run at the speed of a 3B model…
A developer built an open-source alternatives directory using a two-phase ETL pipeline with Turso libSQL and GitHub API. The project required a careful UPSERT strategy to avoid clobbering live star co…
A developer evaluated four free neural TTS options—edge-tts, Kokoro, MeloTTS, and Bark—for use in CI pipelines without a GPU. edge-tts offers broadcast-quality voices via an unofficial Microsoft endpo…
Researchers at BlueDot's Technical AI Safety Project Sprint found that fine-tuning language models on negated documents can still teach them reward-hacking knowledge, leading to emergent misalignment …
Mariana Souza published a tutorial on quantizing and running Meta's Llama 3.2 3B model on Apple Silicon using llama.cpp with Metal GPU acceleration, achieving local inference with Q4_K_M quantization.…
Z.AI's open-weight GLM-5.2 model outperforms GPT-5.5 on the SWE-bench Pro coding benchmark, scoring 62.1 versus 58.6, while costing $1.40 per million input tokens compared to GPT-5.5's $8.00. Released…
Flama 2.0 introduces a CLI tool that allows users to download, package, and serve large language models from HuggingFace with a single command, including a built-in chat interface and production-ready…
An AI/ML student built StudyMate AI, a RAG-based PDF question answering system that uses local embeddings and in-memory vector storage. The project overcame initial retrieval failures by adding pre-ge…
A developer provides a decision tree and code examples to distinguish between RAG (Retrieval-Augmented Generation) and agentic AI architectures. RAG is recommended for answering questions from documen…
NVIDIA has released a GenAI LLM Certification Lab that guides developers through building a production-ready fine-tuning and optimization pipeline. The lab covers data preparation, LoRA fine-tuning wi…
Baidu released Unlimited-OCR, a new model that uses Reference Sliding Window Attention to parse up to 40 PDF pages in a single inference pass, eliminating the linear memory growth of traditional LLM-b…
A developer tested native binary embedding training against post-hoc binarization using a small BERT-mini model and found that native training with a binary loss produced better retrieval results on S…