{"slug": "show-hn-offlineisbetter-efficient-models-for-text-on-cpu", "title": "Show HN: Offlineisbetter: Efficient Models for Text on CPU", "summary": "Offlineisbetter released offline-sentiment-small, a 230M-parameter text model for sentiment analysis that runs on CPU, claiming an F1 of 0.9489 on the stanfordnlp/sst2 validation set at a p95 latency of 80.32 ms on a single thread of a Ryzen 9950X3D. The company said the model, installed via `pip install offlinedemo` and unpacked from a tar archive, is aimed at encoding tasks such as sentiment analysis, text tagging and document retrieval rather than autoregressive generation, and that it outperformed distilbert-base (67M, 66.15 ms, 0.9321 F1), roberta-base (125M, 469.65 ms, 0.9396 F1) and modernbert-base (149M, 530.73 ms, 0.9396 F1) on the same benchmark.", "body_md": "inference for models by *offlineisbetter*.\n\nwe believe that you shouldn't give your data to faceless companies, that you deserve to run text models locally, and that you shouldn't need to buy expensive hardware. so we're building *offlineisbetter*.\n\n*offlineisbetter* models will be for *encoding* tasks: sentiment analysis, text tagging, document retrieval, etc., rather than for *decoding* tasks like autoregressive generation. we believe that it's wasteful and dangerous to depend on cloud apis for frontier language models to do these simple tasks, and it should be almost mindless to download a model to *use* it without dealing with runtimes or quantization formats.\n\ntry our first model yourself. our first model is a small (230m) text model for sentiment analysis called `offline-sentiment-small`.\n\n```\npip install offlinedemo\n```\n\ndownload the model archive from the releases page of this repo and unpack the model.\n\n```\ntar -xvf offline-sentiment-small.tar\n```\n\nrun the demo, passing the inflated directory containing the model checkpoint.\n\n```\nofflinedemo offline-sentiment-small\n```\n\nbelow are benchmarks for `offline-sentiment-small` on binary sentiment classification using the [stanfordnlp/sst2](https://huggingface.co/datasets/stanfordnlp/sst2) validation set. all benchmarks were completed on the ryzen 9950x3d cpu on one thread.\n\n| model | parameters | p95 (ms) | f1 (validation) | \n|---|---|---|---|\n| `offline-sentiment-small` | 230m | 80.32 | 0.9489 | \n| `distilbert-base` | 67m | 66.15 | 0.9321 | \n| `roberta-base` | 125m | 469.65 | 0.9396 | \n| `modernbert-base` | 149m | 530.73 | 0.9396 | \n\nsuccinctly, the core philosophy of *offlineisbetter* is that parameter-efficient and low-latency models should be easily accessible to everybody. of course hugging face and `transformers.pipeline` allows you to run sentiment analysis in three lines of python, but for more parameter-efficient models, already quantized and with optimized computation graphs.", "url": "https://wpnews.pro/news/show-hn-offlineisbetter-efficient-models-for-text-on-cpu", "canonical_source": "https://github.com/offlineisbetter/offlinedemo", "published_at": "2026-09-30 14:42:56+00:00", "updated_at": "2026-09-30 14:50:23.181339+00:00", "lang": "en", "topics": ["natural-language-processing", "machine-learning", "ai-tools", "developer-tools"], "entities": ["Offlineisbetter", "offline-sentiment-small", "stanfordnlp/sst2", "distilbert-base", "roberta-base", "modernbert-base", "Hugging Face", "Ryzen 9950X3D"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-offlineisbetter-efficient-models-for-text-on-cpu", "markdown": "https://wpnews.pro/news/show-hn-offlineisbetter-efficient-models-for-text-on-cpu.md", "text": "https://wpnews.pro/news/show-hn-offlineisbetter-efficient-models-for-text-on-cpu.txt", "jsonld": "https://wpnews.pro/news/show-hn-offlineisbetter-efficient-models-for-text-on-cpu.jsonld"}}