{"slug": "evenlight-i-gave-the-ai-the-smallest-job-in-the-app-and-measured-why", "title": "Evenlight: I gave the AI the smallest job in the app, and measured why", "summary": "A developer built Evenlight, an Android app that captures a single camera photo of the sky, extracts eight horizontal color bands, names the dominant color from a fixed 35-word vocabulary, and prints a card with the date and the sun's altitude. After finding an on-device language model performed poorly at the obvious task, the developer gave it a smaller job, and the app ships with no INTERNET permission and a runtime permission set of exactly {CAMERA}, so measurements stay on the phone. The project's shutter is gated only on an idle flag, with no hour, sun-altitude or scene check, and the developer reports all measurements were taken on an emulator standing in for a phone.", "body_md": "**Tagline:** one balcony, one sky, one card an evening.\n\n**Repo:** [github.com/Burry071/evenlight](https://github.com/Burry071/evenlight) - public, 92 commits, first one\n\n2026-10-08 01:55 +0500, inside the entry window.\n\n**Hacktoberfest 2026, Week 1, theme \"Touch Grass.\"**\n\nEvery entry in this theme seems to be about getting people outside. Mine takes a photo of the sky from a place you\n\nare already sitting and turns it into a card. It does not know whether you went anywhere. I want to be careful about\n\nthat, because it is the most interesting thing about the project: the app claims exactly what it measures and\n\nnothing else.\n\nThe other interesting thing is that I ran a language model on the phone - on an emulator standing in for a phone, for\n\nevery measurement in this post, which is its own story below - found it was bad at the obvious task, and gave it a\n\nsmaller one. That decision is the spine of this post, and every number behind it is either in a file in the\n\nrepo or labelled out loud as a measurement that is not.\n\nThe template asks how this gets people off the screen and into the world, and the honest answer is narrower than\n\nthe question. The app's only input is a photo the camera takes while you are holding it. There is no gallery import\n\nand no way to hand it an evening you did not stand under, so the artefact cannot exist unless somebody went outside\n\nand pointed a phone at the sky. That is the entire mechanism. It is a weak one, and I would rather describe it\n\naccurately than dress it up as a behaviour-change intervention.\n\nI was asked for a sunset notification and did not build one. The reason is arithmetic, not principle: a notification\n\nwants `POST_NOTIFICATIONS`, and the runtime permission set here is exactly `{CAMERA}` with a test asserting it.\n\nEvenlight's one distinctive claim is that it can see nothing but the photo you hand it - there is no `INTERNET`\n\npermission in the manifest, so the measurements cannot leave the phone even if the app wanted them to. Trading a\n\nverifiable fact for a nudge I would not personally obey is a bad deal at any time and a worse one the day before a\n\ndeadline. The argument for the other side is real, though, so here it is: an app that never asks gets opened only\n\nwhen you happen to remember, and the stack in this repo has holes in it for exactly that reason.\n\nOne claim I checked in the code rather than in the design doc, because it is the one that makes this more than a\n\nsunset app: the shutter is gated on nothing but `enabled = !busy`. No hour check, no sun-altitude check, no scene\n\ncheck. \"Evening\" in the UI describes what I use it for, not what it accepts. Point it at a garden at noon, at grass,\n\nat a wall, and it measures eight bands of whatever colour is actually there and names it from the same 35-word\n\nvocabulary. The only hard requirement is that you saved a place and a pair of coordinates once, because the card\n\nprints the sun's altitude - and the app refuses to file a card rather than print an altitude computed for the Gulf\n\nof Guinea.\n\nOne photo, taken fresh by the camera. No gallery import, no presets, no saturation slider.\n\nEight horizontal colour bands, read off the photo. A name, chosen by matching the most colourful band against a\n\nfixed vocabulary of 35 words. A date, and the sun's altitude in degrees at the moment the shutter was pressed. One\n\nline of text.\n\nThen, in the app rather than on the card, every evening you have recorded, oldest first, with a visible gap where an\n\nevening is missing. A streak counter would have hidden the days I did not shoot, and the holes are the honest part\n\nof the data.\n\nThe card is 1080x1440 pixels and the colour on it is the colour that was measured: no filter, no preset, and the\n\nrenderer's only colour input is the eight values. The reason I wrote a second implementation of the sampler in\n\nPython is so that a change to those numbers cannot pass in one language alone.\n\nThe real thing first: three screens off a Pixel 7 on the evening of 10 October, shot from a balcony four minutes\n\nbefore the sunset the phone had computed for 17:49.\n\nShoot is one tap, under a countdown the device computed itself - in 4 min, sunset 17:49, evening 1 - from NOAA\n\nsolar maths and a pair of coordinates typed in once at setup. No location permission, no network to fetch either.\n\nFour minutes before sunset the sun's altitude rounds to 0 degrees, and that is what the card says: not a forecast,\n\nthe altitude at the shutter minute. Eight measured bands, a name the matcher chose from a fixed vocabulary of 35\n\nwords, and one line of wording - here in the template's voice, \"thin light, measured from this sky\", which is what\n\nthe app writes when the model's answer is unusable or absent. The record on the phone says which of the two it\n\nwas; I have not pulled it off the device yet, so I am not going to guess in print.\n\nAnd the stack, in dark mode: one evening, zero holes. The hole is the design - miss an evening and the stack shows\n\nthe gap instead of pretending.\n\nThe five-panel walk below is the emulator's, because setup was never photographed on the phone. It is drawn by\n\n`tools/python/make_screens.py` from five committed screenshots, so it is regenerated rather than collaged, and it\n\nsays on its own face that those panels are emulator panels - which is why the card in its fourth frame reads a sun\n\naltitude of -63 degrees, there being no sun in a synthetic scene. Set that frame beside the phone's 0 degrees and\n\nyou have the two halves of this post in one picture: what the model phrasing was measured on, and what the\n\nmeasurement itself looks like on a real sky.\n\nSetup asks for a place name and a pair of coordinates, once. The tour's other four panels are the countdown and the\n\nthree phone screens above.\n\nTwo cards the app actually filed on the emulator, both reproducible from the JSON beside them:\n\nThe app itself is not a link you can click. It is an APK you build from the repo\n\n(`gradle :app:assembleDebug`, which needs network on a first run for dependencies, then\n\n`adb install -r app/build/outputs/apk/debug/app-debug.apk`),\n\n61,231,992 bytes with no model weights inside it, and the weights are a separate 584 MB gated download whose terms\n\nyou accept in your own browser. That is the one part of the demo I cannot hand you, and it is the reason the two\n\npaths below exist: the measurement is demonstrable without any of it.\n\nI wanted the model to look at eight colours and name the evening. So I ran it on the inputs the app would actually\n\nhand it, before writing any app code. Five designs, a 1B parameter open-weight model on CPU, the same eighteen calls\n\non the last two.\n\n| what I asked for | what came back | \n|---|---|\n| \"name the evening, 2-3 words, no digits\" | obeyed the format, and copied nouns straight out of my input | \n| same, plus a ban list, exactly 2 words, temperature 1.0 | the same echo failure | \n| \"choose from exactly these words\" (the real 35-word vocabulary) | **0/6 on-list.** It answered`sun and stone` ,`sun, concrete` | \n| \"choose one of these three\" | **on-list 17/18** , but**stable across reordering 1/6** | \n| the identical 18 calls on-device, CPU | **on-list 18/18, single-line 18/18, stable 1/6** | \n\nRead the middle row again. Handed the actual vocabulary and told to pick a word from it, it got the answer\n\n*wrong every single time out of six*, inventing phrases that were not on the list. Handed three choices, it was\n\nalmost perfect: 17 of 18 on-list. Then I shuffled the three and it changed its mind in five of six attempts, and the\n\none stable answer reproduced across two independent runtimes - which rules out one runtime's quantisation being the\n\nwhole story, and is the only evidence I have that it is not an artifact of my harness. One case is\n\nworth describing: the same three candidate names in three different orders came back with three different winners.\n\nThat is the whole finding, and it flipped my design. The model is perfectly capable of phrasing something it is\n\nhanded. It is not capable of deciding. So the matcher owns every choice in this app, meaning every choice: which\n\nband wins, which order the bands go in, which word the evening gets. The model writes one line of prose from a word\n\nit did not pick, and it is not allowed to pick anything.\n\n```\nname: brassy\nbands: #1E2438 #2A3350 #3C4A67 #5C6B82 #8A8375 #B08A5E #C2854A #3A322A\nsun: 9 degrees\nWrite one line of four to eight words that contains the word brassy. Do not use any digit and do not add a\ncolour that is not listed.\n```\n\nThe system prompt started as one sentence about what it may do, plus a ban on digits, brackets and newlines, and the\n\nverifier that read the reply had the same shape; both halves turned out wrong and both were rewritten, the prompt into\n\nan allowlist-shaped ask and the gate into the character set below. Then a verifier reads the reply. It rejects the line if the reply is not lowercase, contains a digit, is more than eight words or\n\nfewer than four, or is missing the word it was given. On rejection, the card uses a template line, `brassy light,`, and a counter increments. The card still renders.\n\nmeasured from this sky\n\nThat was the design. Then I put the weights on a device and ran nineteen shots through the installed app, and both\n\nhalves of that paragraph turned out to be wrong about something.\n\nNineteen shots, weights present, engine already warm. **Seven of nineteen** got a line that belonged on the card.\n\nEvery number below is read out of the JSON the app files, and all nineteen records are in the repo at\n\n`docs/data/wording-sample/` with their own recount script beside them, so this is arithmetic you can check rather\n\nthan a claim you have to trust.\n\nThey are two experiments, not one, because the verifier changed halfway through:\n\n| gate | shots | answered | silent | rejected | accepted | accepted wrongly | \n|---|---|---|---|---|---|---|\n| denylist (the first 13) | 13 | 6 | 7 | 1 | 5 | 1 | \n| allowlist (the last 6) | 6 | 6 | 0 | 3 | 3 | 0 | \n| all 19 | 19 | 12 | 7 | 4 | 8 | 1 | \n\nBlending those rows is the dishonest move, and it is tempting. The one wrongly accepted line is a miss the shipped\n\ngate cannot make, so it does not belong in the shipped gate's rate; and a six-shot row cannot carry a rejection\n\nrate at all (3 of 6 rejected reads as \"the new gate is worse\" and means nothing). So: the denylist's\n\nfailure is a finding about denylists, and the allowlist's row is a finding that the sample is too small to publish\n\na rate from.\n\nThat last line - `the **ash** sky glowed softly as dusk fell.` - is the one I would not have found by reading the\n\ncode. The verifier banned `#`, `<`, `{`, `(`. It did\n\nnot ban `*`. The gate is a denylist, which means it has to be right every time, and a 1B model only has to be\n\ncreative once. The card rendered the asterisks, literally, under the name `ash`, and the record said everything was\n\nfine because the line had passed. The denylist became an allowlist: a line may now contain lowercase letters, spaces,\n\nand `. , - '`. An allowlist only has to be right about the punctuation I actually want, and the vocabulary needs\n\nnone beyond that. The cost is honest and I would rather publish it than hide it: any odd character now counts as a\n\nrejection, so the rejection rate goes up.\n\nThe seven silences are the second one, and none of them were visible until the schema could name them. The app filed\n\nall seven as `wording_source: \"none\"`, which in my schema meant \"there is no model on this device.\" There was a model\n\non the device. The field could not tell \"never consulted\" from \"asked and got nothing,\" so any rate built on it was\n\nuncomputable, and it was wrong in the flattering direction: a model that failed looked like a model that was absent.\n\nThere is a fourth value now, `\"no-answer\"`, and that makes the two failures separable instead of one share of\n\nevenings: **did it answer** (12 of 19) and **was that answer valid** (7 of the 12 answers pass today's gate - the\n\nother five are the four the gate rejected plus the markdown line the old gate wrongly passed). One number about\n\nlatency, one about judgement. Both are recountable from the record files, and the script in\n\n`docs/data/wording-sample/PROVENANCE.md` prints the 7.\n\nOnce the seven had their own value, their timestamps told the rest of the story. Every record carries its shutter\n\nminute, and that alone is enough to see it: all seven silences sit in 23:24, 23:25 and 23:26, three minutes that hold\n\nnine shots, of which two answered. The ten shots before and after that window all answered. Same build, same weights,\n\nsame model - the only variable is spacing. I have a sharper version of that sentence from the device's file mtimes at\n\nsecond resolution, that the fast shots went 5-13 s after their neighbour and the answered ones at least 22 s. It is\n\nnot in the repo, because I deleted those records when I put the emulator back the way I found it, so count that half\n\nas my measurement rather than something you can recount.\n\nWhat is in the repo is `docs/data/wording-sample/three-timed-shots.txt`: three shots on the current build, fired\n\nabout a minute apart, each with the time its record first appeared and the time it was written last. The final write\n\nlands between 12.4 s and 24.2 s after the shutter tap, which is 11.3 to 21.4 s after the record already existed. Two\n\nof the three are model lines (24.2 s and 14.3 s after the tap); the 12.4 s one is a rejected line, so the fastest\n\nwrite in the set was the template, not an answer. So a\n\nshot taken twelve seconds later arrives while the previous call is still inside the engine. The seven no-answers are\n\nnot a model that refuses. They are one engine and two callers, and the reason the closing path now holds a guard\n\nacross the whole native call rather than across the timeout.\n\nThe honest half: I fired those shots in a burst because I was testing a code path, not living with an app. One\n\nevening a day is the cadence the thing is for, and in this sample the ten shots that were not fired two-to-four to a\n\nminute all answered. Both numbers are in the table above. The one I would have published without looking is the\n\nflattering one.\n\nTraps I checked before trusting any of it, because a bad sample is worse than no sample:\n\n`SamplerConfig(topK = 20, topP = 0.9, temperature = 0.7)` sets no seed, so I was afraid consecutive\nreplies might come back identical and my \"19 shots\" would really be one. Seven distinct lines out of eight\naccepted, one repeat (`the evening sky held an ash-like glow.` twice). Not pinned, but not independent either.`ash` 13 times and `flint` 6), but the sun altitude read\n`-63°`, `-64°` or `-65°` on every card, because there is no sun in the test scene. So: no claim about weather\nbelongs here, and no claim about evenings either. The one real sky in this post is the phone's, at the top of the\ndemo: one evening, sun 0 degrees, and no weather claim attached to it either.```\nthe stars were bright against an ash-colored\nsky.\n```\n and `the stars shimmered beneath an ash-lit sky.` - in a scene with no stars and a record that measures\nnone. A verifier can only check form, because the only facts this pipeline has are eight colours and one altitude.\nThat is the argument for the whole design: the colour on the card never comes from the model. If I had let the\nmodel choose anything at all, \"stars\" would have been a plausible-looking lie on a card that claims to be a\nmeasurement.\nWhat this is not yet: a timing or a memory number from the phone I own. The phone has filed one real evening - the\n\nthree screens at the top of this post, 10 October, sun 0 degrees - but I have not pulled its record or read its\n\nmemory off the device, so every millisecond and megabyte below is still the emulator's: x86_64 with 3 GB of RAM\n\nand, with the model resident, about 132 MB free and the low-memory killer active. And the emulator's set is\n\nnineteen shots of one synthetic scene, not a season. Two timings, because they are different measurements and\n\naveraging them would be a lie: the bare engine call\n\nin the probe cost 1,080 / 1,247 / 3,502 ms (min / median / max), while inside the app the rewrite lands 11-21 s\n\nafter the record is filed. I have not split that difference up, so I am not going to explain it. Init 4,566 ms, peak\n\nRSS 1,269 MB, both from the emulator.\n\nThe photo is resampled to 64 pixels wide with a hand-written bilinear filter, split into eight bands, and each band\n\nis averaged per channel. Mean luminance uses the sRGB linearisation, so a value at or below 0.03928 divides by 12.92 and\n\nanything above it goes through a power of 2.4. The winning band is the one with the largest gap between its biggest\n\nand smallest channel; a tie goes to the brighter band.\n\nFour details cost real time, and they are all in the code with comments:\n\n**Rounding.** Python's built-in `round` is banker's rounding, so it sends `.5` to the even neighbour. Kotlin's\n\n`Math.round` sends it up. The Python port needed its own `round_half_up`, or the two implementations disagree at\n\nevery half-pixel boundary and the equivalence test fails for a reason neither algorithm caused.\n\n**A clamp that only one fixture can see.** The resample height is derived from the source aspect ratio. When the\n\nsource is narrower than 64 pixels, the derived height exceeds the number of rows the photo actually has, and the\n\nbilinear pass cheerfully interpolates rows that do not exist. Clamped, and one fixture is 32x6 so that no future\n\nchange to that line can pass silently. Every other fixture is at least 64 pixels wide, where the clamp is the\n\nidentity. That is why the odd-sized fixture exists.\n\n**Precision across the border.** One side was doing channel arithmetic in 32-bit floats and the other in 64-bit.\n\nThat disagreement showed up as 13,161 differing `(row, band-pair, channel)` outcomes when the blend formula was\n\nbrute-forced instead of rendered - the kind of thing that looks like a rendering bug until you notice the two sides\n\nround exact `.5` ties opposite ways (row 57, the pair 0 and 47: one language says 0, the other says 1). That sweep\n\nwas a throwaway probe and its script is not in this repository, so this is the one number in the post whose derivation\n\nscript you cannot even look at - the emulator timings above are equally yours to take on trust. What you can check is the fix: `fieldBlend` returns a `Double`, and no cast quantises it on the way to a\n\npixel. Canonical precision is now 64-bit on both sides, and the Python port is told never to touch\n\n`numpy.float32`.\n\n**Name ties.** Two vocabulary words can sit at the same distance from a measured colour. The earlier entry in the\n\nvocabulary wins, and a test asserts that order rather than leaving it to whatever a map iteration does on some\n\nmachine later.\n\nNobody has to install an Android app, and nobody has to download a gigabyte of weights.\n\n```\npython tools/python/skycard.py photo.jpg -o card.png --print-json\npython tools/python/skycard.py --from-record docs/data/wording-sample/2026-10-10-attempt-5.json -o card.png\n```\n\nThe first line prints the eight bands, which band won, and the name, then writes a card. The second redraws the card\n\nthe app filed from the JSON record it wrote, which is how you can check that the picture says what the measurement\n\nsays rather than what I say about it. Timed as whole processes on this box,\n\ntwo runs each: 0.27 s and 0.33 s for a 128x96 fixture, 4.7 s and 6.1 s for a 4032x3024 JPEG at 2 MB (not committed - that second file\n\nis a synthetic gradient with noise in it at the resolution a common phone shoots at, not a photograph - the one real\n\nsky in this repo is the phone's, photographed for the demo above), and the spread between its two runs is why I give\n\nboth numbers instead of a confident\n\none. The band arithmetic is the same arithmetic as the Kotlin, line for line, and a test on both sides asserts\n\nagainst the same five PNG fixtures, so agreement is transitive and neither language can drift without a failure.\n\n89 Kotlin core tests, 49 app tests, 20 Python tests.\n\n`stat` on the emulator's filesystem, not estimated - but they are\nemulator figures, and I have not measured a phone yet, so an ARM cache of a different size is possible.`getprop` and `/proc/meminfo` read off the device, not my memory of this box. The\n10 s give-up marker belongs re-checked there too, because on the emulator it turned out to bound nothing:\n`withTimeoutOrNull` cannot interrupt a blocking native call, and a call that ran 10.5 s still delivered its line\npast the marker.`fontFeatureSettings` is for OpenType features, and `wght` is not one), so the variation is pinned in a\n`<font-family>` XML and loaded through `ResourcesCompat`, because `Resources.getFont` throws on a family file.`CAMERA`, and a test asserts the manifest has no `INTERNET`. Read the next\nsentence before you trust that bullet: the test opens `src/main/AndroidManifest.xml`, the file I wrote, not the\n`assembleDebug`, not a test that fails for you.\n`grep uses-permission -A1 app/build/intermediates/merged_manifests/debug/*/AndroidManifest.xml`.\nKnowing the guarantee is weaker than the test makes it look is the point; a test can hide exactly that from its own\nauthor.\nIt does not measure air quality, pollution, UV, temperature, or your mood. It does not know if you left the\n\nbalcony. It has no reminders, no streak, no cloud sync, and no account. It cannot tell you the sky\n\nlooked better yesterday, because yesterday's photo was taken from a different angle of the same view and the app is\n\nnot that dishonest.\n\nIt does have a `share` link, and I would rather overstate that than let a reader find it later: it is an\n\n`ACTION_SEND` of the card PNG through a `FileProvider`, with a per-URI read-only grant, and the runtime permission\n\nset is still exactly `{CAMERA}`. Nothing leaves the phone by itself, which the manifest test is the guarantee of -\n\nnot my intention.\n\nThere is one real evening in this repo now: 10 October, a balcony, four minutes before a sunset the phone computed\n\nfor 17:49, sun 0 degrees, name thin, filed on a Pixel 7 and photographed at the top of this post. Everything else -\n\nevery record behind the measurements above, and the five PNGs under `fixtures/`, which are drawn, not photographed -\n\nis the emulator's. The rest of the stack is still empty, and it will keep showing that: miss an evening and the hole\n\nstays visible. That is the part of the design that survives my own laziness.\n\nOne thing I did learn, and it is the reason the model ended up with the smallest job in the app: I spent days\n\ntreating the 1/6 as a bug in my own harness. It reproduced across two independent runtimes, which is what the spec\n\nnow says plainly, that this is a model limit and not a quantisation artifact. At some point I stopped debugging and\n\naccepted that the model is genuinely that unreliable at picking. Knowing that is worth more than knowing why my\n\ncode was right, and the only way I found it was to measure before I built.\n\nThree things about this project only work because the pieces are open, and the third one is uncomfortable.\n\nThe model is open-weight - Gemma 3 1B IT, q4, 584,417,280 bytes, run through LiteRT-LM 0.16.1 on the CPU backend -\n\nand that is not a licensing nicety. It is why the app can have no `INTERNET` permission at all. A hosted model puts\n\na network permission in the manifest, and then \"nothing leaves the phone\" becomes a promise about my intentions.\n\nOpen weights turned a privacy claim into a fact about a file, which is the kind of claim a test can hold.\n\nThe arithmetic is open, so the central claim of this post is checkable rather than believable. The band sampler\n\nexists twice - Kotlin on the phone, Python in `tools/python/skycard.py` - and a test on each side asserts against\n\nthe same five PNG fixtures, so neither language can drift without a failure. The nineteen records behind the model's\n\nfailure rate are JSON files in the repo with a recount script beside them. The finding that a 1B model changed its\n\nmind in five of six reorderings of the same three options is a claim about weights you can download and run\n\nyourself. If any of that were closed, you would have to take my word for the number that decided the architecture -\n\nand that number is the only reason the architecture looks the way it does.\n\nThe uncomfortable one: openness is also why I could afford to be pessimistic. Those probe runs cost nothing but\n\ntime, because nobody was metering the calls. Eighteen calls to discover that a model picks badly is a free\n\nexperiment on open weights and an invoice on a hosted API, and an expensive experiment is one you are tempted to\n\nskip and assume your way past instead. Open weights bought the measurement that demoted the model to the smallest\n\njob in the app. I do not think I would have paid to be told I was wrong.\n\nThe repository carries the spec it was built from and the plan that decomposed it, both committed, both written\n\nbefore the code they describe: `docs/superpowers/specs/2026-10-07-evenlight-design.md` and\n\n`docs/superpowers/plans/2026-10-08-evenlight.md`. Both are wrong in places the implementation later corrected, and\n\nsection 12 of the spec is a list of claims from an earlier, discarded design - two of which were simply false and\n\nare named as such. This was built with AI agents over four days. Leaving the arguments in the repo is the only\n\nhonest way to show that; a summary of the process would be a story about it instead.\n\nCounts at the commit this post describes: **89 tests in `:core`, 49 in `:app`, 20 in Python**, no skips and no\n\nfailures. The debug APK is 61,231,992 bytes. The weights are not committed - GitHub's hard per-file limit is\n\n100 MB and the file is 584,417,280 bytes, so `models/GET_MODEL.md` ships the URL, filename, byte size and SHA-256\n\ninstead, and `*.litertlm` is gitignored so the omission cannot be an accident.\n\n**Overall.**\n\n**Best Use of Gemma.** Gemma 3 1B IT, q4, on-device through LiteRT-LM 0.16.1, CPU backend, no network. I want to be\n\nexact about the size of the role, because the honest version is smaller than the category name suggests: the model\n\nwrites one line of prose per card and chooses nothing. Every decision - which band wins, what order the eight go in,\n\nwhich of 35 words the evening gets - belongs to a deterministic matcher, because I measured the model failing at\n\nexactly that job before I wrote the app. If \"best use\" means the largest role for the model, this is not that entry.\n\nIf it means the most measured one, the eighteen probe calls and nineteen device shots are in the sections above, and\n\nthe records are in the repo.\n\nNo other partner category. Nothing here is hosted, and I would rather leave a category empty than stretch a claim\n\nto fill it.", "url": "https://wpnews.pro/news/evenlight-i-gave-the-ai-the-smallest-job-in-the-app-and-measured-why", "canonical_source": "https://dev.to/fahadyy/evenlight-i-gave-the-ai-the-smallest-job-in-the-app-and-measured-why-4d8o", "published_at": "2026-10-11 11:17:10+00:00", "updated_at": "2026-10-11 11:22:08.847248+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "ai-products"], "entities": ["Evenlight", "Pixel 7", "GitHub", "Hacktoberfest"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/evenlight-i-gave-the-ai-the-smallest-job-in-the-app-and-measured-why", "markdown": "https://wpnews.pro/news/evenlight-i-gave-the-ai-the-smallest-job-in-the-app-and-measured-why.md", "text": "https://wpnews.pro/news/evenlight-i-gave-the-ai-the-smallest-job-in-the-app-and-measured-why.txt", "jsonld": "https://wpnews.pro/news/evenlight-i-gave-the-ai-the-smallest-job-in-the-app-and-measured-why.jsonld"}}