cd/entity/Ollama· home› entities› Ollama
grep -l @ollama /news/*.json | wc -l → 1521

Ollama

mentions 1521 type Organization page 15/77 feed RSS

// recent coverage 1521 mentions

22:05
2026-09-12
promptcube3.com
large-language-models

Gemini Forum, Qwen Coder local setup

A developer resolved CUDA out-of-memory crashes running Qwen2.5-Coder-32B on a 24GB RTX 3090 by capping the context window at 8k tokens instead of 32k via a custom Ollama Modelfile, cutting latency fr…

15:45
2026-09-12
github.com
ai-agents

Orcrist: A Coding Agent using LLM state machines

Orcrist is a new domain-specific language for state machines whose states are executed by an LLM, paired with a desktop coding agent that runs on it. Before touching a task, the agent writes an Orcris…

11:42
2026-09-12
dev.to
large-language-models

ROCm vs Vulkan for AMD Local LLM Hosting: 2026 Guide

A 2026 guide compares ROCm and Vulkan as backends for hosting local LLMs on AMD GPUs, concluding that the choice depends on the inference engine, GPU generation, and workload rather than being interch…

18:07
2026-09-11
dev.to
ai-infrastructure

Mohdel 1.0: a self-hosted LLM gateway and SDK for Node

A developer released Mohdel 1.0.0, an MIT-licensed self-hosted LLM gateway and SDK for Node that unifies access to Anthropic, OpenAI, Gemini, Mistral, Groq, xAI, Cerebras, Fireworks, DeepSeek, Qwen Cl…

18:06
2026-09-11
dev.to
large-language-models

Why LLM Load Tests Are Costing You Thousands

A developer detailed how load testing LLM API integrations can cost thousands of dollars in metered tokens, citing a $3,000 bill from a failed 100,000-request test. The engineer, who built an Autonomo…

15:53
2026-09-11
blog.kilo.ai
large-language-models

How to Choose a Local LLM: Models, Hardware, and Quantization

Atomic Chat published a guest blog guide on selecting local large language models, recommending that users start with a GGUF Q4_K_M quantization if it fits their hardware. The guide provides memory-es…

12:20
2026-09-11
github.com
ai-agents

Pizza Bot: a local-first inbox for long-running AI agents

Amazon released Pizza Bot, an open-source, local-first inbox for long-running AI agents, under the Apache 2.0 license. The tool runs on a stateful DeepAgents/LangGraph runtime with an Electron desktop…

11:49
2026-09-11
ory.com
ai-tools

Make Claude Code Faster and Cheaper with Ory Lumen

Ory released Ory Lumen, an open-source local semantic code search engine that runs as an MCP server alongside Claude Code and cut Claude Code runtime by up to 53% and API costs by up to 39% in SWE-ben…

← prev page 15 / 77 next →
// co-occurs with top 8 entities
// topics top 6 topics