Jev: State In, Typed Decisions Out
TypeSafe announced Jev, a model that takes unstructured state as input and returns typed, probabilistic decisions — Choice, Score, and Noul — rather than free text, claiming it is 20-200x faster and 4…
TypeSafe announced Jev, a model that takes unstructured state as input and returns typed, probabilistic decisions — Choice, Score, and Noul — rather than free text, claiming it is 20-200x faster and 4…
A community derivative of Qwen2.5-1.5B-Instruct, processed with Heretic to reduce refusal behavior, is available as a Q4_K_M GGUF file on Hugging Face from user saidutta69. The derivative is intended …
Signal Drift, a CRT-noir hacking thriller, ships with locally run AI characters powered by Qwen2.5-1.5B-Instruct, a 1.54-billion-parameter model fine-tuned with a custom LoRA and quantized to Q4_K_M, …
An engineer investigating speculative decoding speedup found that the algorithm's performance ceiling is determined by acceptance rate and cost ratio, not just hardware overhead. Testing on Apple Sili…
A developer known as flirp spent 21 GPU-hours training Qwen2.5-1.5B-Instruct to reason in latent space without decoding tokens, but the model's hidden-state reasoning collapsed into content-free attra…
An engineer built NilaMind, an open-source Android app that runs a 1.5B-parameter language model entirely on-device for mental health conversations, with no cloud, account, or analytics. The app uses …
A new open-source tool, jlens-gguf, implements Anthropic's Jacobian Lens for GGUF models, enabling interactive visualization and live steering of language model activations directly through a web UI. …
A developer tested an aggressive version of latent-space reasoning on a 1.5B-parameter language model, where the model pauses during generation to run parallel hidden-state rollouts without decoding t…
UmarTransit-1B, the first open-source large language model fine-tuned for public transit systems and GTFS data, has been released. Built by fine-tuning Qwen2.5-1.5B-Instruct using QLoRA, the model spe…
Qwen released Qwen2.5-1.5B-Instruct, an Apache-2.0 licensed 1.5B parameter instruct model for local chat and agent tasks, requiring 8-16 GB RAM/VRAM. The model is available on Hugging Bay with externa…
A developer built an LLM-as-a-judge from scratch using Qwen2.5-1.5B-Instruct and tested it against the LMSYS Chatbot Arena dataset with human votes. The judge scored answers independently and agreed w…
A developer benchmarked speculative decoding using Qwen2.5-0.5B-Instruct as the draft model and Qwen2.5-1.5B-Instruct as the target model on a CPU. Across code, JSON, and story generation tasks, specu…