cd/entity/Ollama· home entities Ollama
grep -l @ollama /news/*.json | wc -l → 989

Ollama

mentions 989 type Organization page 4/50 feed RSS

// recent coverage 989 mentions

21:52
2026-08-14
discuss.huggingface.co
artificial-intelligence

Looking for help building an offline, local AI desktop app

A developer is seeking paid help to build a privacy-focused desktop app for Windows and Mac that runs an open-source AI model fully offline, with no cloud, account, or telemetry, and requires a custom…

17:45
2026-08-14
promptcube3.com
artificial-intelligence

Cutting my AI subscription bill by 60% was surprisingly easy

A user reports cutting their AI subscription bill by 60% by replacing paid writing assistants, PDF analyzers, meeting note tools, image generators, and copywriting services with local open-source mode…

17:32
2026-08-14
promptcube3.com
large-language-models

Why is Qwen 2.5-27B acting so erratic on my local setup?

A developer reports that Qwen 2.5-27B, a large language model, exhibits erratic behavior and degraded output quality on a local setup with an NVIDIA RTX 3090 (24GB VRAM), particularly when the context…

16:45
2026-08-14
promptcube3.com
artificial-intelligence

vLLM beats Ollama by 20x once you hit high concurrency

VLLM outperforms Ollama by nearly 20x in throughput at high concurrency, according to benchmark tests running Llama 3.1 8B on an NVIDIA A100 40GB, with vLLM peaking at 793 tokens per second versus Oll…

16:35
2026-08-14
opencrawling.org
ai-infrastructure

The Qdrant Output Connector

Qdrant, the Rust-based vector database, has been integrated into OpenCrawling's event-driven microservice architecture via a new Qdrant Output Connector that enables sub-millisecond ACL pre-filtering …

12:16
2026-08-14
dev.to
large-language-models

vLLM vs Ollama: Production Serving 2026

A developer's comparison of vLLM and Ollama for LLM serving in 2026 shows that while Ollama is simpler for single-user local use, vLLM outperforms it dramatically under concurrency, with throughput up…

07:42
2026-08-14
pipe-lang.com
artificial-intelligence

📚 RAG in ~10 Lines — No Vector Database, No pip install

Pipe, a pipeline-native language, now supports retrieval-augmented generation (RAG) in about 10 lines of code without a vector database, framework, or external dependencies, according to a blog post i…

07:39
2026-08-14
github.com
ai-infrastructure

Show HN: Miser – Cost-Optimised AI Gateway

Miser, an open-source Rust-based AI gateway, routes OpenAI-compatible requests to the cheapest capable model via OpenRouter, achieving 92.0% exact tier accuracy with sub-millisecond latency on a 25-ca…

03:15
2026-08-14
github.blog
ai-tools

GitHub Copilot weekly releases — August 10

GitHub's weekly Copilot update (August 10) introduces Kimi K3 to all paid plans, makes Agent Plugins 1.0 generally available, and adds MAI-Code-1.1-Flash with image understanding. The update also brin…

14:25
2026-08-13
promptcube3.com
developer-tools

Stop wasting hours on manual PR reviews with these AI setups

A developer shares AI setups to streamline code reviews, including using the Continue extension with a local Ollama instance for autocomplete and Claude 3.5 Sonnet for chat, and a 'Reviewer Persona' p…

14:00
2026-08-13
kdnuggets.com
artificial-intelligence

Building a Streaming Local AI Agent

A new tutorial by an unnamed author demonstrates building a streaming local AI agent that monitors Wikipedia's live edit feed for vandalism using Ollama, with a two-stage funnel to filter events befor…

← prev page 4 / 50 next →
// co-occurs with top 8 entities
// topics top 6 topics