ChimpBench – Chimpanzee Population Simulator ChimpBench, a personal experiment in using decision models to drive behavior, simulates communities of wild chimpanzees in a 3D rainforest modeled on Kibale National Park in western Uganda, with each chimp's choices decided by rules or a small AI model running on the user's computer and compared against field data from wild chimps. The simulation reproduces documented wild behaviors including silent border patrols, with males at Ngogo setting out about every 10 days (Watts & Mitani 2001), and lethal territorial raids that let Ngogo, an unusually large community, grow its territory over a decade (Mitani, Watts & Amsler 2010). The forest grows several of the food tree species recorded at Ngogo and Kanyawara, which sit only 12 km apart yet differ noticeably in food trees (Potts, Chapman & Lwanga 2009). A chimpanzee society simulation, based on wild chimps. ChimpBench simulates communities of wild chimpanzees in a 3D rainforest modeled on Kibale, Uganda. Each chimp feeds, rests, travels, grooms, fights and reconciles, patrols, hunts, raises young and ages, through day and night and the seasons. A personal experiment in using decision models to drive behavior. Each chimp’s choices come from what it senses and how it feels, and are decided by rules or by a small AI model running on your computer. The results are compared with field data from wild chimps. Dawn, day one. The field log and range map on the left, the forest in the middle, and the communities on the right, with the selected community's parties and members below. Colors mark the communities: West, East and North. The forest is modeled on Kibale National Park in western Uganda, home to Ngogo and Kanyawara, two wild chimpanzee communities studied for decades. The individual chimps and their histories are made up. How they behave comes from what field researchers there, and at other sites, have published. Unevenly matched communities One community starts with more grown males than its neighbours. That's the kind of edge Ngogo, an unusually large community, has over its neighbors, and it grew its territory after a decade of lethal raids Mitani, Watts & Amsler 2010 . Kibale's trees Ngogo and Kanyawara sit only 12 km apart yet grow noticeably different food trees Potts, Chapman & Lwanga 2009 . The simulated forest grows several of the species they recorded, giant figs included. Kibale's weather and light Two rainy seasons, afternoon thunderstorms, cool dawns and warm afternoons, and a real equatorial sun and moon. A stream The chimps drink from it and cross it only at shallow fords. Wild behaviours in the simulation Border patrols are silent Males walk the edge of their territory in single file, quiet, listening. At Ngogo one set out about every 10 days Watts & Mitani 2001 . They compare numbers before a fight Play a stranger's call from a hidden speaker and males move in when their side has the numbers Wilson, Hauser & Wrangham 2001 . Most fights are bluffs Hair on end, charging, dragging branches, drumming on roots. It looks scary, and that's the point: it rarely comes to blows. They make up Opponents often reconcile soon after a fight, and bystanders comfort the loser Kutsukake & Castles 2004 . Daughters move away As teenagers, most females join a neighboring community. Males stay where they were born, for life. A new nest every night At dusk every chimp past toddler age bends branches into a new nest high in a tree on how they pick nest trees: Samson & Hunt 2014 . Males hunt together Parties with more males hunt red colobus monkeys more often and catch more Mitani & Watts 1999 , then share the meat with allies 2001 . They warn others about snakes At a snake, chimps call more when nearby group members haven't seen it yet Crockford et al. 2012 . They remember old groupmates Apes recognize former groupmates decades later, most of all the ones they got along with Lewis et al. 2023 . A long childhood Babies nurse for almost five years Bray et al. 2017 , and a mother has a new one about every five Wallis 1997 . Individual chimps Each chimp is a full individual: a body with needs, a mood, a personality, a rank, a family, friendships, grudges and a memory. The inspector shows all of it. Below is an example card for one chimp, Koruza, from one run. Close view. Following Tavuni, West's alpha, as he mate-guards while others forage, play and pant-grunt around him. His Overview opens with one line: what he's doing, and his mood. Press C in the app for this view. Keep-clear off. The same paused moment with the fade switched off a debug switch : a trunk hides part of the party. Keep-clear on. The close camera fades whatever stands between you and the chimps you're following. Koruza West communityAdult male, 24Rank 3 of 8 males Right now: grooming Mbelo Body Needs rise and fall all day. Food fixes hunger, grooming fixes loneliness, sleep fixes tiredness. Hunger Thirst Energy Lonely Stress Health Mood and personality Mood comes from what's happening. Personality is set at birth and sticks. calmexcitedplayfulfearfulaggressivedistressed Bold Sociable Aggressive Playful Skills They grow with practice: play builds climbing, feeding builds foraging, hunts build hunting. Climbing Foraging Hunting Social Rank Males climb by winning contests. The one on top is the alpha, the boss of the community. 1Tavuni 2Mbelo 3Koruzayou are here 4Sanaki 5+ 4 more Family Mom and siblings matter most. Dad is in the records, but no chimp knows who he is. Friends and rivals Grooming and backing each other up build bonds. A fight leaves tension that fades over a few weeks, faster if they make up. Mbeloally Rukasobrother Sanakitense, fought yesterday bondtension Memory A diary that summarizes itself. Fresh moments stay detailed, months get condensed, and each year becomes one entry kept for life. Recent momentsthe latest few Just now: started grooming Mbelo 2 h ago: lost a fight with Sanaki 5 h ago: heard strangers to the east Monthly summariesthe past year Last month: groomed most with Mbelo; backed by Rukaso 6 times; made up with Sanaki twice Yearly summarieskept for life Last year: climbed to rank 3; closest friend Mbelo When a model decides for a chimp, it gets a small slice of this: how the chimp feels, who's around, and a few memory lines about them. Daily routine The simulation keeps real equatorial time: sunrise and sunset just before seven, about twelve hours of daylight all year. The timeline shows what most chimps are doing across a typical day. Light daylight Calls Dawn chorus06:24–07:24 · everyone joins in with pant-hoots, their loud rising “hoo-hoo-HOO” that carries far Dusk chorus18:00–19:00 · one more round before bed Doing Asleepuntil about 06:50 · still in last night's nests Breakfast, mostly fruit06:50–11:30 · feeding at fruit trees and moving between them Siesta, grooming11:30–14:30 · resting through the heat, grooming friends More feeding14:30–17:50 · back to the trees Build a nest, sleepfrom about 17:50 · a fresh nest, then sleep Patrols A border patrol may set out08:00–15:30 · every week or two, a group of males heads for the border Storms Thunderstorms most likely12:30–19:00 · about two-thirds of the rain falls in the afternoon 0608101214161820 Rainy seasons Kibale gets rain from March to May and again from September to November, with drier spells in between. Fruit follows the rain about six weeks later. Figs are different: each tree fruits on its own schedule, so some figs are usually ripe when little else is. 15 °Cat dawn 24 °Cafter lunch Rain per month, mm, as simulatedwet season 62J 78F 150M 185A 160M 78J 60J 110A 155S 185O 180N 95D About 1,650 mm a year. Every new forest opens at 06:30 on 28 September, in the middle of the rains. Night. The chimps asleep in fresh nests, grouped into nesting parties. A storm. Most of the party hunches down to wait it out, and Kidoti did a rain display see the log . Tavuni, still reacting to a stranger's call, patrols. Decisions Nothing is scripted. Whenever a chimp finishes what it was doing, or something interrupts it a charge, a storm, a stranger's call , it stops and makes a fresh decision. This repeats for every chimp, all day. Sees and hearsonly what's nearby, no peeking at the map Remembersfriends, fights, fruit trees, water Lists optionsonly moves that are possible right now Picks onethe built-in rules or a decision model Actseats, grooms, charges, naps… World changesbellies, bonds, ranks, everyone nearby Every chimp, all daythen back to step 1 …and back to step 1. Rules The built-in rules score each option with a simple recipe hungry plus fruit nearby means go eat and take the top one. They run every chimp the decision model isn't running. Async default The forest keeps moving while the decision model thinks. If an answer comes back late, the rules fill in and the chimp carries on. Lockstep The forest pauses at every model decision and waits for the answer, so every pick lands exactly when it was asked. Outside the app, test runs hand the same chimps to other deciders: trained adapters, the stand-ins that imitate them in long runs, Jev, and rules that keep their plan. GLiNER2.5-Decide In the app, the decisions the rules don't make come from fastino/GLiNER2.5-Decide, a small classifier that picks from a list of options it's given, running locally. It is one of several deciders ChimpBench tests: the built-in rules, three trained adapters for GLiNER baseline, aggressive and collaborative , small stand-in networks that imitate those adapters in multi-year runs, and Jev TypeSafe System One as an offline test arm. The free part of the decisive test is done; the paid Jev part is approved and pending. What it is Its model card calls it a specialist classifier for operational decisions: you give it some text and a set of labels at run time, and it scores the labels. It doesn't chat, reason out loud or answer open questions. That's exactly the shape of a chimp's next move: one situation, a short menu, pick one. Why small and local A chimp decides every few minutes of forest time, and dozens of them can be deciding at once. The model has to keep up, run on a laptop with no account, no network and no bill, and only ever choose, never invent. A big chat model would be slower and could make up moves that don't exist. Model fastino/GLiNER2.5-Decide Size 340M parameters Built on DeBERTa-v3-large encoder License Apache 2.0 Runs on Apple silicon GPU MPS , fp16 Per decision ~500 input tokens, 0.25–0.4 s Options per call 2 to 8 Input limit 1,280 tokens, never truncated observe A private snapshot of what this chimp sees, hears, feels and remembers. Nothing it couldn't know. Legal options Up to 8 moves that are possible right now, each written as a plain sentence, always including the rules' pick. GLiNER scores The model scores every option against the snapshot, in about 0.4 s. Check The server validates the answer; the simulation re-checks the move is still legal. Act and trace The chimp acts. The trace keeps every score and what the rules would have picked. Late answers In async mode the forest doesn't wait. A late answer is checked against the newer state and applied only if the move is still legal; after six forest minutes the rules decide instead. Lockstep The clock stops on the exact tick a model-driven chimp starts waiting and resumes when the answer lands. After 6 real seconds the rules take over, so it can't hang. Trained adapters and stand-ins Three LoRA adapters, trained on labelled decisions, change which option GLiNER picks: a field-expert baseline, an aggressive temperament and a collaborative one. They run in headless test runs, not in the app. Stand-ins, small networks fitted to imitate each adapter, run the five-year tests at rules speed. Offline test arms Jev runs only in offline tests. The free arms rules, rules that keep their plan, a utility formula and random choice have run on five fresh worlds. The paid Jev arms are approved with a $10 cap and haven't run yet. What the model actually read418 tokens · 478 ms meTavuni, adult male, 27 y, West community; the alpha, rank 1 of 8 males; mood calm; timid; solitary; playful feelingmild hunger now07:19 dawn; thunderstorm, 16 °C; party of 5 with 2 adult males nearbySanaki: my ally, adult male, ranks below me, 2 m, close bond · Yerubi: juvenile female, 2 m · 2 more memoriesJust now: a heavy storm broke eventsA heavy thunderstorm is pouring down “You are a field primatologist. Choose what this wild eastern chimpanzee would most plausibly do next, given only what it perceives, feels and remembers…” Sit hunched through the storm AI pickrules pick57% Groom my ally Sanaki to keep his support17% Groom Yerubi, juvenile7% Play with the young Yerubi6% Sit and rest nearby4% Forage on leaves and pith nearby4% Charging display to assert my alpha status3% Travel toward Rukaso's pant-hoots3% The Mind tab. Tavuni is at the edge of his range with one companion. GLiNER picks rest 62% where the rules would have foraged, and the history shows it disagreeing with the rules now and then. Every row keeps the rules' score next to the model's. What the first real runs showed The untuned model's first 3,546 logged decisions It matches words It scores an option largely by how its text relates to the situation's text. So each need and each option is phrased in the same words "food", "water" , and only active needs are named: a cue that's always there nudges every decision the same way. Hungry, before7% Hungry, after60% Thirst was its weak spot The same wording barely moved thirst: very thirsty chimps almost never picked the stream, even when it was offered. That led to the fix below. Thirsty, before0.3% Thirsty, after7% Strangers trigger a response When a chimp had just heard strangers, real or from the playback experiment, most picks became a response: flee, call back, patrol toward them, display or charge. All quiet9% Strangers heard69% It nests at dusk Between about 17:50 and 19:30, with a nest on the menu, it built one more than half the time. At midday a nest is hardly ever even offered. Dusk55% It agrees with the rules about half the time 48% across that log, and 40–65% across runs. When they differ it leans social: grooming was its most common pick. Same pick48% Repeated facts bias it Repeating a partner's details made it favor that partner, and "currently mate-guarding" next to a "Mate-guard…" option locked it into continuing 7 of 36 contexts at 90%+, versus 2 without the echo . Now each fact appears once, and ids never reach the model. The thirst fix Two changes. The state now names every urgent need, not just the strongest a very thirsty chimp that was slightly hungrier used to hear only "hunger" . And when a bodily need is urgent, the matching option repeats it in the state's own words: "water, eases severe thirst, needed now". 120 real contexts from the log were replayed through the model before and after the fix. When thirst is the stronger need, 10 of 12 chimps now drink; the misses are hunger winning. Moderate thirst still rarely drinks, but the rules rarely drink at that level either. 120 replayed real contextsbeforeafter the fix Very thirsty → drink30% → 50% Hungry → forage75% → 94% Strangers → respond63% → 75% Dusk → nest50% → 63% Storm → shelter56% → 56% The replayed set has its own before numbers, so they differ from the log above. Society Social events are not scripted. Rank, alliances, conflict and reconciliation, territories, hunting, family life and parties each come from a small set of rules that run for every chimp, every day. Where a rule follows a field study, the study is cited. Rank and alphas Males rise by winning contests, scored with Elo ratings Neumann et al. 2011 , and lower-ranked chimps greet higher ones with a pant-grunt, a breathy "you're the boss". The alpha usually falls to a rival with friends. Females queue by age and tenure instead Foerster et al. 2016 . Alliances Chimps who groom each other, back each other up and share meat become allies. When a friend gets charged, allies may jump in, and that's how a smaller male can topple a stronger one. Fights and making up Most conflicts are charges and bluffs; about 1 in 20 gets physical. Afterwards, rivals often reconcile and friends console the loser Kutsukake & Castles 2004 . A fight that never gets patched up leaves tension that fades over a few weeks, modeled on how chimps' relationships vary in value and compatibility Fraser, Schino & Aureli 2008 . Territories and patrols Every week or two, a group of males patrols the border in silence Watts & Mitani 2001 . Meeting strangers, they count heads: outnumber them and they charge; outnumbered, they slip quietly home. Most encounters are only heard, not seen Wilson et al. 2012 . Hunting and sharing When several males spot red colobus monkeys, a hunt can start. More hunters, better odds, as at Ngogo Mitani & Watts 1999 . The catch gets shared with friends, allies and family who beg. Babies and families A baby rides on mom's belly, then her back, and nurses until 4 or 5. The next one comes about five years after the last Wallis 1997 . Death rates by age follow Ngogo's Wood et al. 2017 , and orphans are sometimes adopted by an older sibling. Females change communities Around age 12, most young females move to a neighboring community and start at the bottom of the female pecking order. Males stay home for life. Food and weather A community splits into parties, temporary subgroups that merge and split all day: big ones at a fig feast, small ones when fruit is scarce. Storms send everyone to shelter. Kinship. Every matriline as a tree; dashed lines trace genetic sires. Dominance. Male and female ladders by Elo score, per community. Press T in the app. Field experiments Field experiments test how chimps react when they perceive or experience a specific event, such as a stranger's call played from the edge of their range or a model snake on a trail. In the app, the Experiments panel places the same kinds of events in the simulated forest, and the Mind tab shows a chimp's choice just before and just after. Stranger call playback A hidden speaker plays a strange male's pant-hoot. Do they go look, call back or sneak away? Wilson, Hauser & Wrangham 2001 Snake model A fake viper on the trail. Who spots it, and do they warn the ones who haven't? Crockford et al. 2012 Fig feast A giant fig ripens all at once. Watch everyone show up. Storm A heavy downpour, right now. Everyone hunkers down; a few males do a rain display. Drought A few days with hardly any fruit. Parties shrink and squabbles over food go up. Remove the alpha The boss vanishes. Watch the struggle to fill the top spot. Monkeys arrive A group of red colobus moves in overhead. Hunting chance. Watts & Mitani 2002 Before and after The Mind tab puts the chimp's choice right before the experiment next to the one right after. Stranger playback, before and after. In this example run, Tavuni was foraging 40% from the app's decision model, GLiNER2.5-Decide . After a stranger's call about 18 m away, the model picks flee at 100% and the rules agree, while Sanaki and Rukaso set off on patrol. Field playbacks found the same pattern: small parties back off, bigger male parties move in Wilson, Hauser & Wrangham 2001 . Architecture One rule organizes the code: only the simulation changes the world. The 3D view, the sound and the interface read it; the decision loop asks it for a chimp's view and returns a checked choice. Decision model: which models are being tested? The engine Simulation Every chimp's body, mind and relationships, plus the forest they live in. The only part that changes the world. Clock Moves the world forward in fixed 15-second steps. Saves Your simulations, stored in your browser. Reload and carry on. The brains Decision loop Picks which chimps ask the decision model, sends it what that chimp sees, and hands a checked pick back to the simulation. Decision modelReads one chimp's view and returns one pick; it never touches the world directly. Several models are being tested.Which models What you see and hear 3D forest Terrain, trees, sky, rain and every chimp, animated. Read only. Sound Calls, birds, rain and thunder, placed around your camera. Read only. Interface Panels, family trees, rank ladders, experiments and the Mind tab. It sends your controls speed, experiments to the simulation. Decision models being tested Rules The built-in baseline: a simple recipe scores each option. GLiNER2.5-Decide The small local model that runs live in the app. Trained GLiNER versions Three temperaments, baseline, aggressive and collaborative, tested outside the app. Jev TypeSafe Tested offline. The free test is done; the paid test is pending. Made with TypeScript, three.js for the 3D, Web Audio for sound, and SQLite running in the browser for saves. Chimp data The field data used to recreate a realistic forest and realistic chimps: open datasets from long-term studies at Ngogo, Gombe and Taï set the forest and its fruit seasons, and give the results the simulated chimps are checked against. The field data usedNgogoGombeTaï Eight open datasets from three sites, 1978 to 2024. One builds the simulated forest; the other seven test the simulated chimps. Loading the datasets… Each dataset is used under its licence CC0 or CC BY 4.0 and credited wherever its numbers appear. Raw records never leave data/raw/: this guide shows counts, shares, rates and normalized maps only. Every target, pass or fail Each square is one field target. Held-out targets were never used for tuning, and most of the scored ones still fail. Loading the scorecard… Targets and pass bands: data/targets.json, from 124 cited studies. Verdicts: the latest year-long field-profile runs of the virtual researcher. A tuned, encoded or compromised target never counts as a pass. Behaviour by behaviourfield bandeach simulated seedoutside the bandM, F: scored per sex Grooming time, fruit in the diet and male bonds land in the wild bands on the tuning seeds, less clearly on fresh ones. Being improved: hunting success, patrol rate and home range are well outside their bands. Loading… Each row has its own scale, from zero. Bands, sites and sources: data/targets.json; values: the tuning seeds. Fruit calendar Ripe fruit is chimps' main food, and its timing affects party size and travel. At Ngogo, researchers recorded ripe fruit on the same marked trees every month from 1998 to 2017 Potts et al. 2020 . The simulated forest field profile uses this record directly: its fruit follows the Ngogo months in order, year by year. Twenty years of ripe fruit at Ngogo Ripe fruit comes in pulses, and no two years repeat. The simulated forest steps through these years in order, month by month. Loading the phenology record… Data: Ngogo, Kibale National Park, 1998–2017 Potts et al. 2020, Biotropica; data Dryad, CC0 . The ripe-fruit score is the source's crop-weighted index. Monthly aggregates are derived by ChimpBench; not endorsed by the authors. Ripe fruit by species and month The eight Ngogo species in the simulated forest fruit at different times, and the figs carry a little fruit almost all year, which keeps some food available when other species have none. Loading the species records… Share of monitored stems with ripe fruit, per species and month same source, CC0 . The ninth species, Ficus sansibarica, isn't on the Ngogo transect and uses the fig average. Movement Three open datasets record where wild chimps went: GPS fixes from Ngogo 2011–2023 , 30-minute focal records from Taï 2013–2016 and 15-minute focal follows from Gombe 2000–2003 . The same measures are computed for real and simulated chimps: home-range area and core area 95% and 50% kernels , how use spreads across the range, 30-minute step length, turning angle and how straight a day's path is. The virtual field observer samples the simulated chimps on the same schedules. Status: the way use spreads within a range matches Ngogo. Being improved: ranges are about three times too small for the group size, and day paths double back too often. Similar Loading the comparison… Different Not comparable yet Range area: the simulated communities' ranges are at least 2.7× smaller than the band for wild communities of about 22 members 1.8 km² against 5–16 km² , while the way use spreads inside a range matches. Day paths: net distance divided by distance walked is …, against … at Taï and … at Gombe, and simulated chimps reverse direction more often. Parties: about 3 chimps, against about 6 at Taï. Being improved. These charts come from an earlier run. Later changes brought turning and party size closer to Taï, but not straightness 0.23 against 0.50 : each change that raised straightness made turning or party size worse. Movement work is paused until the next full test run, and the charts will be redrawn after it. Simulated ranges, to scalesimulatedwildtop of the size-matched band Against wild communities of the same size, simulated ranges are several times smaller being improved . Ngogo is shown for scale only: it has about nine times as many chimps, so it is not a like-for-like comparison. Loading… Size-matched band: data/targets.json T-RNG-1 Kanyawara, Budongo Sonso and Waibira; Taï . Taï groups of 11–23 members, yearly 95% kernels Lemoine et al. 2020, CC BY 4.0 . Ngogo community kernel, 2011–2014 Sandel et al. 2026, CC BY 4.0 . Simulated: field profile, 95% kernels, five seeds. Each chimp · NgogoCommunity · NgogoTerritory · Taï Why the maps look abstract. Chimp location data can help poachers, so it's handled carefully and this guide never shows where any chimp actually was. Each map is centred on its range, scaled by its own range radius r , rotated to its long axis and mirrored. Only the shape is left, and shape is what's being compared. Six full days: Gombe vs simulatedDifferent Measured in range radii, a simulated day covers more than twice the distance of a Gombe day but ends about as far from its start, because it doubles back more often being improved . Loading the paths… Six full days per side at evenly spaced straightness ranks, so neither side is cherry-picked. Real: Gombe National Park, Kasekela community, 2000–2003, 30-minute records from 15-minute focal follows Pusey & Schroepfer-Walker 2013; data Dryad, CC0 . Paths are re-drawn from the records after privacy normalization; not endorsed by the authors. Distance covered in 30 minutes … Taï, western chimpanzeesGombe, eastern chimpanzeesSimulateddotted: median Turning between 30-minute stepsDifferent At both wild sites, most turns between 30-minute steps are small. Simulated chimps make more sharp reversals being improved . How much the ranges of Ngogo's two halves overlap, year by year 1 = the same range, 0 = none . The community polarized in 2015 and was two groups by 2018. No simulated community has split yet, so the simulated lines stay flat. Ngogo, West vs CentralSimulated: best split inside one communitySimulated: two neighbouring communities How the day is spentshare of 30-minute records The daily budget is close to Taï's: a bit more feeding, a bit less rest. Loading… Taï only records rest, travel and feeding, so grooming and other social time count as rest here. A boundary patrol is a group of males walking the edge of their range, stopping to listen and sometimes crossing into a neighbour’s range. Patrol records from Ngogo, Gombe and Taï give how often patrols happen, who joins and how long they last, and the simulated patrols are measured the same way. Loading the comparison… Patrols by month of the yearPer monthShare of the year Gombe is the fair seasonal reference: its patrols come from daily follows all year round, and they happen in every month. Ngogo's monthly counts mostly show when observers were in the field. Loading the patrol records… Each site's patrols by calendar month, averaged over its study years. Gombe 1978–2007 Massaro et al. 2022; data Dryad, CC0 . Ngogo 1996–2015 Langergraber et al. 2017; data Dryad, CC0 . Monthly counts only; no dates or names. Patrols by the numbers Measure Ngogo Gombe Band Simulated Loading… Numbers computed from open records say open data; the rest are published figures. Each band was fixed before any run of the corrected patrol model. Ngogo's community is far bigger than the simulated ones 24–44 males , so shares and rates are the fair comparison, not head counts. Who joins a patrol In a small community most males go on most patrols. In Ngogo's huge one, the typical male joins about a third. Loading… One dot per male, anonymous. Gombe Massaro et al. 2022, CC0 : share of observed patrol opportunities joined over the whole study. Ngogo Langergraber et al. 2017, CC0 : share of the patrols a male was listed on that he joined, males listed on 20 or more. Advance or retreat after a border stop … Taï, western chimpanzeesSimulatedshaded: the band Where border stops happen … Taï border stops Before and after the Ngogo split Territory gained after killings Patrols vs ordinary visits to the edge Gombe, 1978–2007: patrols against ordinary visits to the edge of the range, seen on the same daily follows. Measures not yet comparable Only summaries leave the raw files. The patrol records name individual chimps and give exact dates. This guide publishes counts, shares, rates and distances measured in range radii, never a name, a place or a day. ChimpBench can hand the same chimps to different decision policies and score each one with the same virtual field researcher, against the same wild bands. So far, the best result comes from the simplest change: the built-in rules, told to keep their plan and to choose with some randomness. Rules The built-in recipe that runs the app. It was tuned on travel share, party size and day range, which gives it a home advantage. Even so, in the free test its chimps travel too much, groom too little and walk too far each day. Untuned GLiNER2.5-Decide The model as downloaded. It leans social: over three observed days its chimps groomed 28% of daylight wild: 8–18% and walked 0.15 km a day wild: 1.5–3.5 km . When aggression was on offer it chose it about five times as often as the field-expert labels 0.29 against 0.06 . Trained adapters Three small add-ons to GLiNER LoRA adapters , trained on the same 900 decisions with different labels: a field-expert baseline, an aggressive temperament and a collaborative one. The labels come from an AI labeller following a written rubric, not from field data. On 200 held-out decisions, each adapter matches its own labels far better than the untuned model 0.74, 0.72 and 0.81 against 0.40–0.54 , and the temperaments separate clearly. Jev TypeSafe System One A hosted model that turns a described situation and a set of options into probabilities. Driving one community at a time for three days, it chose less aggression and more grooming than rules, and picked social options 64% of the time when they were offered expert labels: 36% . A paid decisive test is approved but hasn't run. Rules + plan The same rules with two changes. A chimp keeps its current plan until something important changes, such as a need crossing a threshold, the time of day moving on, an interruption, the plan ending, or 90 minutes passing. And it picks from the rules' scores with some randomness instead of always taking the top one. In the free test this cut the gap to wild chimps by 55%, on all five worlds. Two controls: a formula and random choice The utility formula scores each option by how much hunger, thirst, tiredness and loneliness it relieves per hour, walk included, using the simulation's own rates. No model is involved. Random choice picks any legal option. Random choice beat rules on distance, but its chimps went hungry, so it doesn't count. What the model runs showed Real models: 3-day runs. Five-year runs: stand-ins. Hungry chimps eat leaves instead of walking to fruit This is the adapters' main failure. At strong or severe hunger, with leaves the only food nearby and a remembered fruit tree on the menu, they eat leaves in place 97–100% of the time and never make the trip. Rules make it 14–27% of the time. Leaves feed about half as fast as ripe fruit, so a chimp living on leaves has to feed about 75% of the day. Share that walk to the fruit, at strong hunger: Rules14% Adapters0% They feed too much and travel too little Over three observed days the adapters' chimps fed 59–71% of daylight wild: 33–50% . Over five simulated years, run with stand-ins that imitate each adapter, the trained communities fed 71–82% of daylight and travelled 2–9% wild: 12–25% . Adding randomness can't fix this for the adapters: they put nearly all their weight on one option, so it shifts their choices by 2 points or less. The temperaments are real The aggressive adapter charged about four times as often as rules and was the only one to cause major injuries. The baseline and collaborative adapters made no charges at all, which is probably less aggressive than wild males. The collaborative one groomed about 3.5 times as often as rules. Charges per adult per day: Rules2.1 Aggressive8.7 Untuned3.9 Baseline0 Collaborative0 Jev0.1 Jev and untuned GLiNER barely travel Jev picked a trip in 0.2% of the decisions where one was offered. Picking in proportion to its probabilities, instead of always taking the top one, would still send its chimps travelling in only about 1.5% of decisions. Untuned GLiNER's chimps walked 0.15 km a day. Three observed days with the real modelsfield map, 2 seeds rules: 3 Measure Wild band Rules Untuned Baseline Aggressive Collaborative Feedingshare of daylight 33–50% 46%in band 41%in band 71%above 59%above 62%above Travelshare of daylight 12–25% 26%above 7%below 13%in band 16%in band 9%below Groomingshare of daylight 8–18% 7%below 28%above 7%below 10%in band 20%above Male day rangekm walked per day 1.5–3.5 km 3.6 kmabove 0.15 kmbelow 1.2 kmbelow 2.1 kmin band 2.2 kmin band Party sizechimps together 3–9 2.8below 3.3in band 2.8below 2.05below 4.5in band Means over the seeds, after a 180-day rules burn-in. Seed ranges are wide aggressive grooming ran from 4% to 16% , so read directions, not ranks. These runs used an older copy of the code than the free test below, so compare within a table, not across. Free test: distance from the wild bandslower is better Keeping the plan and adding some randomness cut the gap to wild chimps by 55%, and beat rules on all five worlds. The paid Jev arms haven't run yet. Rules2.04 Utility formula1.83 Random choice1.37 Rules + plan0.93 Jev as shipped: pendingJev with facts and randomness: pendingSame, facts shuffled: pending The distance sums nine rows: feeding, travel and grooming for each sex, rest, party size and male day range. A row counts 0 inside the wild band; otherwise it counts the gap to the nearest edge, divided by the band width. Mean of five fresh worlds seeds 6501–6905 , all three communities on one policy, five scored days after a 180-day rules burn-in, measured on simulation truth. Pre-registered before any run, after a four-judge review of two Jev designs: docs/staging/jev-decisive-test.md. Results: artifacts/decide-ft/jev-test/free-arms.md. Lactating females are left out of the score until a known energy problem in the simulation is fixed they go hungry even under rules , and are reported separately. The free test, policy by policy5 worlds each Policy Distance Better than rules Median hunger Result Rulesas shipped 2.04 — 0.64 Control Rules + plankeeps its plan, picks with some randomness 0.93 5 of 5 worlds 0.66 Best so far Utility formulaneeds relieved per hour, no model 1.83 3 of 5 0.39 Small gain Random choiceany legal option 1.37 5 of 5 0.86 Non-viable Jev as shippedcurrent packet, top pick pending pending pending Pending Jev with facts and randomnessfood values and walking times in the options pending pending pending Pending Same, facts shuffled1 world: does Jev read the facts? pending pending pending Pending Hunger runs from 0 full to 1, for all adults. A policy is non-viable if its median adult hunger is more than 0.10 above rules', or its lactating females' median reaches 0.95; random choice fails both 0.86 and 0.97 . Rules miss on travel, grooming and male day range. Rules + plan brings travel and day range into the band, but grooming stays low and parties stay a little small. No training on Jev's answers. TypeSafe's customer agreement §2.3 b forbids using Jev's outputs to train a model that imitates it. So Jev can't be distilled into a fast stand-in for long runs without TypeSafe's written permission, and its answers carry a do-not-train marker. Still pending The paid Jev test. Approved with a hard $10 cap, not run yet: it waits for a spending guard tested against a fake server. Jev counts as helping only if its version with facts and randomness beats rules by at least 0.10 on average and on 4 of 5 worlds, beats both rules + plan and the utility formula by at least 0.05 on 4 of 5 worlds, and keeps its chimps fed. Rules + plan as the default policy. Approved and pre-registered as stage C13. It can be switched off, and the combined proof run will judge it. That run hasn't happened yet. Long runs of the trained adapters. The five-year numbers come from stand-ins: small networks fitted to imitate each adapter, which agree with it on 67–77% of decisions in worlds they never saw. They are not the adapters themselves. Credits The tools, decision models, sound and field data behind ChimpBench, and the studies it draws on, each with its licence, are listed on the Credits page.