Evenlight: I gave the AI the smallest job in the app, and measured why A developer built Evenlight, an Android app that captures a single camera photo of the sky, extracts eight horizontal color bands, names the dominant color from a fixed 35-word vocabulary, and prints a card with the date and the sun's altitude. After finding an on-device language model performed poorly at the obvious task, the developer gave it a smaller job, and the app ships with no INTERNET permission and a runtime permission set of exactly {CAMERA}, so measurements stay on the phone. The project's shutter is gated only on an idle flag, with no hour, sun-altitude or scene check, and the developer reports all measurements were taken on an emulator standing in for a phone. Tagline: one balcony, one sky, one card an evening. Repo: github.com/Burry071/evenlight https://github.com/Burry071/evenlight - public, 92 commits, first one 2026-10-08 01:55 +0500, inside the entry window. Hacktoberfest 2026, Week 1, theme "Touch Grass." Every entry in this theme seems to be about getting people outside. Mine takes a photo of the sky from a place you are already sitting and turns it into a card. It does not know whether you went anywhere. I want to be careful about that, because it is the most interesting thing about the project: the app claims exactly what it measures and nothing else. The other interesting thing is that I ran a language model on the phone - on an emulator standing in for a phone, for every measurement in this post, which is its own story below - found it was bad at the obvious task, and gave it a smaller one. That decision is the spine of this post, and every number behind it is either in a file in the repo or labelled out loud as a measurement that is not. The template asks how this gets people off the screen and into the world, and the honest answer is narrower than the question. The app's only input is a photo the camera takes while you are holding it. There is no gallery import and no way to hand it an evening you did not stand under, so the artefact cannot exist unless somebody went outside and pointed a phone at the sky. That is the entire mechanism. It is a weak one, and I would rather describe it accurately than dress it up as a behaviour-change intervention. I was asked for a sunset notification and did not build one. The reason is arithmetic, not principle: a notification wants POST NOTIFICATIONS , and the runtime permission set here is exactly {CAMERA} with a test asserting it. Evenlight's one distinctive claim is that it can see nothing but the photo you hand it - there is no INTERNET permission in the manifest, so the measurements cannot leave the phone even if the app wanted them to. Trading a verifiable fact for a nudge I would not personally obey is a bad deal at any time and a worse one the day before a deadline. The argument for the other side is real, though, so here it is: an app that never asks gets opened only when you happen to remember, and the stack in this repo has holes in it for exactly that reason. One claim I checked in the code rather than in the design doc, because it is the one that makes this more than a sunset app: the shutter is gated on nothing but enabled = busy . No hour check, no sun-altitude check, no scene check. "Evening" in the UI describes what I use it for, not what it accepts. Point it at a garden at noon, at grass, at a wall, and it measures eight bands of whatever colour is actually there and names it from the same 35-word vocabulary. The only hard requirement is that you saved a place and a pair of coordinates once, because the card prints the sun's altitude - and the app refuses to file a card rather than print an altitude computed for the Gulf of Guinea. One photo, taken fresh by the camera. No gallery import, no presets, no saturation slider. Eight horizontal colour bands, read off the photo. A name, chosen by matching the most colourful band against a fixed vocabulary of 35 words. A date, and the sun's altitude in degrees at the moment the shutter was pressed. One line of text. Then, in the app rather than on the card, every evening you have recorded, oldest first, with a visible gap where an evening is missing. A streak counter would have hidden the days I did not shoot, and the holes are the honest part of the data. The card is 1080x1440 pixels and the colour on it is the colour that was measured: no filter, no preset, and the renderer's only colour input is the eight values. The reason I wrote a second implementation of the sampler in Python is so that a change to those numbers cannot pass in one language alone. The real thing first: three screens off a Pixel 7 on the evening of 10 October, shot from a balcony four minutes before the sunset the phone had computed for 17:49. Shoot is one tap, under a countdown the device computed itself - in 4 min, sunset 17:49, evening 1 - from NOAA solar maths and a pair of coordinates typed in once at setup. No location permission, no network to fetch either. Four minutes before sunset the sun's altitude rounds to 0 degrees, and that is what the card says: not a forecast, the altitude at the shutter minute. Eight measured bands, a name the matcher chose from a fixed vocabulary of 35 words, and one line of wording - here in the template's voice, "thin light, measured from this sky", which is what the app writes when the model's answer is unusable or absent. The record on the phone says which of the two it was; I have not pulled it off the device yet, so I am not going to guess in print. And the stack, in dark mode: one evening, zero holes. The hole is the design - miss an evening and the stack shows the gap instead of pretending. The five-panel walk below is the emulator's, because setup was never photographed on the phone. It is drawn by tools/python/make screens.py from five committed screenshots, so it is regenerated rather than collaged, and it says on its own face that those panels are emulator panels - which is why the card in its fourth frame reads a sun altitude of -63 degrees, there being no sun in a synthetic scene. Set that frame beside the phone's 0 degrees and you have the two halves of this post in one picture: what the model phrasing was measured on, and what the measurement itself looks like on a real sky. Setup asks for a place name and a pair of coordinates, once. The tour's other four panels are the countdown and the three phone screens above. Two cards the app actually filed on the emulator, both reproducible from the JSON beside them: The app itself is not a link you can click. It is an APK you build from the repo gradle :app:assembleDebug , which needs network on a first run for dependencies, then adb install -r app/build/outputs/apk/debug/app-debug.apk , 61,231,992 bytes with no model weights inside it, and the weights are a separate 584 MB gated download whose terms you accept in your own browser. That is the one part of the demo I cannot hand you, and it is the reason the two paths below exist: the measurement is demonstrable without any of it. I wanted the model to look at eight colours and name the evening. So I ran it on the inputs the app would actually hand it, before writing any app code. Five designs, a 1B parameter open-weight model on CPU, the same eighteen calls on the last two. | what I asked for | what came back | |---|---| | "name the evening, 2-3 words, no digits" | obeyed the format, and copied nouns straight out of my input | | same, plus a ban list, exactly 2 words, temperature 1.0 | the same echo failure | | "choose from exactly these words" the real 35-word vocabulary | 0/6 on-list. It answered sun and stone , sun, concrete | | "choose one of these three" | on-list 17/18 , but stable across reordering 1/6 | | the identical 18 calls on-device, CPU | on-list 18/18, single-line 18/18, stable 1/6 | Read the middle row again. Handed the actual vocabulary and told to pick a word from it, it got the answer wrong every single time out of six , inventing phrases that were not on the list. Handed three choices, it was almost perfect: 17 of 18 on-list. Then I shuffled the three and it changed its mind in five of six attempts, and the one stable answer reproduced across two independent runtimes - which rules out one runtime's quantisation being the whole story, and is the only evidence I have that it is not an artifact of my harness. One case is worth describing: the same three candidate names in three different orders came back with three different winners. That is the whole finding, and it flipped my design. The model is perfectly capable of phrasing something it is handed. It is not capable of deciding. So the matcher owns every choice in this app, meaning every choice: which band wins, which order the bands go in, which word the evening gets. The model writes one line of prose from a word it did not pick, and it is not allowed to pick anything. name: brassy bands: 1E2438 2A3350 3C4A67 5C6B82 8A8375 B08A5E C2854A 3A322A sun: 9 degrees Write one line of four to eight words that contains the word brassy. Do not use any digit and do not add a colour that is not listed. The system prompt started as one sentence about what it may do, plus a ban on digits, brackets and newlines, and the verifier that read the reply had the same shape; both halves turned out wrong and both were rewritten, the prompt into an allowlist-shaped ask and the gate into the character set below. Then a verifier reads the reply. It rejects the line if the reply is not lowercase, contains a digit, is more than eight words or fewer than four, or is missing the word it was given. On rejection, the card uses a template line, brassy light, , and a counter increments. The card still renders. measured from this sky That was the design. Then I put the weights on a device and ran nineteen shots through the installed app, and both halves of that paragraph turned out to be wrong about something. Nineteen shots, weights present, engine already warm. Seven of nineteen got a line that belonged on the card. Every number below is read out of the JSON the app files, and all nineteen records are in the repo at docs/data/wording-sample/ with their own recount script beside them, so this is arithmetic you can check rather than a claim you have to trust. They are two experiments, not one, because the verifier changed halfway through: | gate | shots | answered | silent | rejected | accepted | accepted wrongly | |---|---|---|---|---|---|---| | denylist the first 13 | 13 | 6 | 7 | 1 | 5 | 1 | | allowlist the last 6 | 6 | 6 | 0 | 3 | 3 | 0 | | all 19 | 19 | 12 | 7 | 4 | 8 | 1 | Blending those rows is the dishonest move, and it is tempting. The one wrongly accepted line is a miss the shipped gate cannot make, so it does not belong in the shipped gate's rate; and a six-shot row cannot carry a rejection rate at all 3 of 6 rejected reads as "the new gate is worse" and means nothing . So: the denylist's failure is a finding about denylists, and the allowlist's row is a finding that the sample is too small to publish a rate from. That last line - the ash sky glowed softly as dusk fell. - is the one I would not have found by reading the code. The verifier banned , < , { , . It did not ban . The gate is a denylist, which means it has to be right every time, and a 1B model only has to be creative once. The card rendered the asterisks, literally, under the name ash , and the record said everything was fine because the line had passed. The denylist became an allowlist: a line may now contain lowercase letters, spaces, and . , - ' . An allowlist only has to be right about the punctuation I actually want, and the vocabulary needs none beyond that. The cost is honest and I would rather publish it than hide it: any odd character now counts as a rejection, so the rejection rate goes up. The seven silences are the second one, and none of them were visible until the schema could name them. The app filed all seven as wording source: "none" , which in my schema meant "there is no model on this device." There was a model on the device. The field could not tell "never consulted" from "asked and got nothing," so any rate built on it was uncomputable, and it was wrong in the flattering direction: a model that failed looked like a model that was absent. There is a fourth value now, "no-answer" , and that makes the two failures separable instead of one share of evenings: did it answer 12 of 19 and was that answer valid 7 of the 12 answers pass today's gate - the other five are the four the gate rejected plus the markdown line the old gate wrongly passed . One number about latency, one about judgement. Both are recountable from the record files, and the script in docs/data/wording-sample/PROVENANCE.md prints the 7. Once the seven had their own value, their timestamps told the rest of the story. Every record carries its shutter minute, and that alone is enough to see it: all seven silences sit in 23:24, 23:25 and 23:26, three minutes that hold nine shots, of which two answered. The ten shots before and after that window all answered. Same build, same weights, same model - the only variable is spacing. I have a sharper version of that sentence from the device's file mtimes at second resolution, that the fast shots went 5-13 s after their neighbour and the answered ones at least 22 s. It is not in the repo, because I deleted those records when I put the emulator back the way I found it, so count that half as my measurement rather than something you can recount. What is in the repo is docs/data/wording-sample/three-timed-shots.txt : three shots on the current build, fired about a minute apart, each with the time its record first appeared and the time it was written last. The final write lands between 12.4 s and 24.2 s after the shutter tap, which is 11.3 to 21.4 s after the record already existed. Two of the three are model lines 24.2 s and 14.3 s after the tap ; the 12.4 s one is a rejected line, so the fastest write in the set was the template, not an answer. So a shot taken twelve seconds later arrives while the previous call is still inside the engine. The seven no-answers are not a model that refuses. They are one engine and two callers, and the reason the closing path now holds a guard across the whole native call rather than across the timeout. The honest half: I fired those shots in a burst because I was testing a code path, not living with an app. One evening a day is the cadence the thing is for, and in this sample the ten shots that were not fired two-to-four to a minute all answered. Both numbers are in the table above. The one I would have published without looking is the flattering one. Traps I checked before trusting any of it, because a bad sample is worse than no sample: SamplerConfig topK = 20, topP = 0.9, temperature = 0.7 sets no seed, so I was afraid consecutive replies might come back identical and my "19 shots" would really be one. Seven distinct lines out of eight accepted, one repeat the evening sky held an ash-like glow. twice . Not pinned, but not independent either. ash 13 times and flint 6 , but the sun altitude read -63° , -64° or -65° on every card, because there is no sun in the test scene. So: no claim about weather belongs here, and no claim about evenings either. The one real sky in this post is the phone's, at the top of the demo: one evening, sun 0 degrees, and no weather claim attached to it either. the stars were bright against an ash-colored sky. and the stars shimmered beneath an ash-lit sky. - in a scene with no stars and a record that measures none. A verifier can only check form, because the only facts this pipeline has are eight colours and one altitude. That is the argument for the whole design: the colour on the card never comes from the model. If I had let the model choose anything at all, "stars" would have been a plausible-looking lie on a card that claims to be a measurement. What this is not yet: a timing or a memory number from the phone I own. The phone has filed one real evening - the three screens at the top of this post, 10 October, sun 0 degrees - but I have not pulled its record or read its memory off the device, so every millisecond and megabyte below is still the emulator's: x86 64 with 3 GB of RAM and, with the model resident, about 132 MB free and the low-memory killer active. And the emulator's set is nineteen shots of one synthetic scene, not a season. Two timings, because they are different measurements and averaging them would be a lie: the bare engine call in the probe cost 1,080 / 1,247 / 3,502 ms min / median / max , while inside the app the rewrite lands 11-21 s after the record is filed. I have not split that difference up, so I am not going to explain it. Init 4,566 ms, peak RSS 1,269 MB, both from the emulator. The photo is resampled to 64 pixels wide with a hand-written bilinear filter, split into eight bands, and each band is averaged per channel. Mean luminance uses the sRGB linearisation, so a value at or below 0.03928 divides by 12.92 and anything above it goes through a power of 2.4. The winning band is the one with the largest gap between its biggest and smallest channel; a tie goes to the brighter band. Four details cost real time, and they are all in the code with comments: Rounding. Python's built-in round is banker's rounding, so it sends .5 to the even neighbour. Kotlin's Math.round sends it up. The Python port needed its own round half up , or the two implementations disagree at every half-pixel boundary and the equivalence test fails for a reason neither algorithm caused. A clamp that only one fixture can see. The resample height is derived from the source aspect ratio. When the source is narrower than 64 pixels, the derived height exceeds the number of rows the photo actually has, and the bilinear pass cheerfully interpolates rows that do not exist. Clamped, and one fixture is 32x6 so that no future change to that line can pass silently. Every other fixture is at least 64 pixels wide, where the clamp is the identity. That is why the odd-sized fixture exists. Precision across the border. One side was doing channel arithmetic in 32-bit floats and the other in 64-bit. That disagreement showed up as 13,161 differing row, band-pair, channel outcomes when the blend formula was brute-forced instead of rendered - the kind of thing that looks like a rendering bug until you notice the two sides round exact .5 ties opposite ways row 57, the pair 0 and 47: one language says 0, the other says 1 . That sweep was a throwaway probe and its script is not in this repository, so this is the one number in the post whose derivation script you cannot even look at - the emulator timings above are equally yours to take on trust. What you can check is the fix: fieldBlend returns a Double , and no cast quantises it on the way to a pixel. Canonical precision is now 64-bit on both sides, and the Python port is told never to touch numpy.float32 . Name ties. Two vocabulary words can sit at the same distance from a measured colour. The earlier entry in the vocabulary wins, and a test asserts that order rather than leaving it to whatever a map iteration does on some machine later. Nobody has to install an Android app, and nobody has to download a gigabyte of weights. python tools/python/skycard.py photo.jpg -o card.png --print-json python tools/python/skycard.py --from-record docs/data/wording-sample/2026-10-10-attempt-5.json -o card.png The first line prints the eight bands, which band won, and the name, then writes a card. The second redraws the card the app filed from the JSON record it wrote, which is how you can check that the picture says what the measurement says rather than what I say about it. Timed as whole processes on this box, two runs each: 0.27 s and 0.33 s for a 128x96 fixture, 4.7 s and 6.1 s for a 4032x3024 JPEG at 2 MB not committed - that second file is a synthetic gradient with noise in it at the resolution a common phone shoots at, not a photograph - the one real sky in this repo is the phone's, photographed for the demo above , and the spread between its two runs is why I give both numbers instead of a confident one. The band arithmetic is the same arithmetic as the Kotlin, line for line, and a test on both sides asserts against the same five PNG fixtures, so agreement is transitive and neither language can drift without a failure. 89 Kotlin core tests, 49 app tests, 20 Python tests. stat on the emulator's filesystem, not estimated - but they are emulator figures, and I have not measured a phone yet, so an ARM cache of a different size is possible. getprop and /proc/meminfo read off the device, not my memory of this box. The 10 s give-up marker belongs re-checked there too, because on the emulator it turned out to bound nothing: withTimeoutOrNull cannot interrupt a blocking native call, and a call that ran 10.5 s still delivered its line past the marker. fontFeatureSettings is for OpenType features, and wght is not one , so the variation is pinned in a