cd/entity/Ollama· home› entities› Ollama
grep -l @ollama /news/*.json | wc -l → 1547

Ollama

mentions 1547 type Organization page 32/78 feed RSS

// recent coverage 1547 mentions

21:52
2026-08-14
discuss.huggingface.co
artificial-intelligence

Looking for help building an offline, local AI desktop app

A developer is seeking paid help to build a privacy-focused desktop app for Windows and Mac that runs an open-source AI model fully offline, with no cloud, account, or telemetry, and requires a custom…

17:45
2026-08-14
promptcube3.com
artificial-intelligence

Cutting my AI subscription bill by 60% was surprisingly easy

A user reports cutting their AI subscription bill by 60% by replacing paid writing assistants, PDF analyzers, meeting note tools, image generators, and copywriting services with local open-source mode…

17:32
2026-08-14
promptcube3.com
large-language-models

Why is Qwen 2.5-27B acting so erratic on my local setup?

A developer reports that Qwen 2.5-27B, a large language model, exhibits erratic behavior and degraded output quality on a local setup with an NVIDIA RTX 3090 (24GB VRAM), particularly when the context…

16:45
2026-08-14
promptcube3.com
artificial-intelligence

vLLM beats Ollama by 20x once you hit high concurrency

VLLM outperforms Ollama by nearly 20x in throughput at high concurrency, according to benchmark tests running Llama 3.1 8B on an NVIDIA A100 40GB, with vLLM peaking at 793 tokens per second versus Oll…

16:35
2026-08-14
opencrawling.org
ai-infrastructure

The Qdrant Output Connector

Qdrant, the Rust-based vector database, has been integrated into OpenCrawling's event-driven microservice architecture via a new Qdrant Output Connector that enables sub-millisecond ACL pre-filtering …

12:16
2026-08-14
dev.to
large-language-models

vLLM vs Ollama: Production Serving 2026

A developer's comparison of vLLM and Ollama for LLM serving in 2026 shows that while Ollama is simpler for single-user local use, vLLM outperforms it dramatically under concurrency, with throughput up…

07:42
2026-08-14
pipe-lang.com
artificial-intelligence

📚 RAG in ~10 Lines — No Vector Database, No pip install

Pipe, a pipeline-native language, now supports retrieval-augmented generation (RAG) in about 10 lines of code without a vector database, framework, or external dependencies, according to a blog post i…

07:39
2026-08-14
github.com
ai-infrastructure

Show HN: Miser – Cost-Optimised AI Gateway

Miser, an open-source Rust-based AI gateway, routes OpenAI-compatible requests to the cheapest capable model via OpenRouter, achieving 92.0% exact tier accuracy with sub-millisecond latency on a 25-ca…

03:15
2026-08-14
github.blog
ai-tools

GitHub Copilot weekly releases — August 10

GitHub's weekly Copilot update (August 10) introduces Kimi K3 to all paid plans, makes Agent Plugins 1.0 generally available, and adds MAI-Code-1.1-Flash with image understanding. The update also brin…

00:00
2026-08-14
pipelineandprompts.com
large-language-models

Compared llama3.2:1b vs llama3.2:3b Memory Footprint

Ollama's llama3.2:3b model consumes 2.5 GB of memory (2.47 GB RSS) versus 1.5 GB (1.24 GB RSS) for the 1b model, showing that tripling parameters roughly doubles memory footprint. The test also reveal…

14:25
2026-08-13
promptcube3.com
developer-tools

Stop wasting hours on manual PR reviews with these AI setups

A developer shares AI setups to streamline code reviews, including using the Continue extension with a local Ollama instance for autocomplete and Claude 3.5 Sonnet for chat, and a 'Reviewer Persona' p…

← prev page 32 / 78 next →
// co-occurs with top 8 entities
// topics top 6 topics