{"slug": "chimpbench-chimpanzee-population-simulator", "title": "ChimpBench – Chimpanzee Population Simulator", "summary": "ChimpBench, a personal experiment in using decision models to drive behavior, simulates communities of wild chimpanzees in a 3D rainforest modeled on Kibale National Park in western Uganda, with each chimp's choices decided by rules or a small AI model running on the user's computer and compared against field data from wild chimps. The simulation reproduces documented wild behaviors including silent border patrols, with males at Ngogo setting out about every 10 days (Watts & Mitani 2001), and lethal territorial raids that let Ngogo, an unusually large community, grow its territory over a decade (Mitani, Watts & Amsler 2010). The forest grows several of the food tree species recorded at Ngogo and Kanyawara, which sit only 12 km apart yet differ noticeably in food trees (Potts, Chapman & Lwanga 2009).", "body_md": "A chimpanzee society simulation, based on wild chimps.\n\nChimpBench simulates communities of wild chimpanzees in a 3D rainforest modeled on Kibale, Uganda. Each chimp feeds, rests, travels, grooms, fights and reconciles, patrols, hunts, raises young and ages, through day and night and the seasons.\n\nA personal experiment in using decision models to drive behavior. Each chimp’s choices come from what it senses and how it feels, and are decided by rules or by a small AI model running on your computer. The results are compared with field data from wild chimps.\n\nDawn, day one. The field log and range map on the left, the forest in the middle, and the communities on the right, with the selected community's parties and members below. Colors mark the communities: West, East and North.\n\nThe forest is modeled on Kibale National Park in western Uganda, home to Ngogo and Kanyawara, two wild chimpanzee communities studied for decades. The individual chimps and their histories are made up. How they behave comes from what field researchers there, and at other sites, have published.\n\nUnevenly matched communities\n\nOne community starts with more grown males than its neighbours. That's the kind of edge Ngogo, an unusually large community, has over its neighbors, and it grew its territory after a decade of lethal raids (Mitani, Watts & Amsler 2010).\n\nKibale's trees\n\nNgogo and Kanyawara sit only 12 km apart yet grow noticeably different food trees (Potts, Chapman & Lwanga 2009). The simulated forest grows several of the species they recorded, giant figs included.\n\nKibale's weather and light\n\nTwo rainy seasons, afternoon thunderstorms, cool dawns and warm afternoons, and a real equatorial sun and moon.\n\nA stream\n\nThe chimps drink from it and cross it only at shallow fords.\n\nWild behaviours in the simulation\n\nBorder patrols are silent\n\nMales walk the edge of their territory in single file, quiet, listening. At Ngogo one set out about every 10 days (Watts & Mitani 2001).\n\nThey compare numbers before a fight\n\nPlay a stranger's call from a hidden speaker and males move in when their side has the numbers (Wilson, Hauser & Wrangham 2001).\n\nMost fights are bluffs\n\nHair on end, charging, dragging branches, drumming on roots. It looks scary, and that's the point: it rarely comes to blows.\n\nThey make up\n\nOpponents often reconcile soon after a fight, and bystanders comfort the loser (Kutsukake & Castles 2004).\n\nDaughters move away\n\nAs teenagers, most females join a neighboring community. Males stay where they were born, for life.\n\nA new nest every night\n\nAt dusk every chimp past toddler age bends branches into a new nest high in a tree (on how they pick nest trees: Samson & Hunt 2014).\n\nMales hunt together\n\nParties with more males hunt red colobus monkeys more often and catch more (Mitani & Watts 1999), then share the meat with allies (2001).\n\nThey warn others about snakes\n\nAt a snake, chimps call more when nearby group members haven't seen it yet (Crockford et al. 2012).\n\nThey remember old groupmates\n\nApes recognize former groupmates decades later, most of all the ones they got along with (Lewis et al. 2023).\n\nA long childhood\n\nBabies nurse for almost five years (Bray et al. 2017), and a mother has a new one about every five (Wallis 1997).\n\nIndividual chimps\n\nEach chimp is a full individual: a body with needs, a mood, a personality, a rank, a family, friendships, grudges and a memory. The inspector shows all of it. Below is an example card for one chimp, Koruza, from one run.\n\nClose view. Following Tavuni, West's alpha, as he mate-guards while others forage, play and pant-grunt around him. His Overview opens with one line: what he's doing, and his mood. Press C in the app for this view.\n\nKeep-clear off. The same paused moment with the fade switched off (a debug switch): a trunk hides part of the party.\n\nKeep-clear on. The close camera fades whatever stands between you and the chimps you're following.\n\nKoruza\n\nWest communityAdult male, 24Rank 3 of 8 males\n\nRight now: grooming Mbelo\n\nBody\n\nNeeds rise and fall all day. Food fixes hunger, grooming fixes loneliness, sleep fixes tiredness.\n\nHunger\n\nThirst\n\nEnergy\n\nLonely\n\nStress\n\nHealth\n\nMood and personality\n\nMood comes from what's happening. Personality is set at birth and sticks.\n\ncalmexcitedplayfulfearfulaggressivedistressed\n\nBold\n\nSociable\n\nAggressive\n\nPlayful\n\nSkills\n\nThey grow with practice: play builds climbing, feeding builds foraging, hunts build hunting.\n\nClimbing\n\nForaging\n\nHunting\n\nSocial\n\nRank\n\nMales climb by winning contests. The one on top is the alpha, the boss of the community.\n\n1Tavuni\n\n2Mbelo\n\n3Koruzayou are here\n\n4Sanaki\n\n5+ 4 more\n\nFamily\n\nMom and siblings matter most. Dad is in the records, but no chimp knows who he is.\n\nFriends and rivals\n\nGrooming and backing each other up build bonds. A fight leaves tension that fades over a few weeks, faster if they make up.\n\nMbeloally\n\nRukasobrother\n\nSanakitense, fought yesterday\n\nbondtension\n\nMemory\n\nA diary that summarizes itself. Fresh moments stay detailed, months get condensed, and each year becomes one entry kept for life.\n\nRecent momentsthe latest few\n\nJust now: started grooming Mbelo\n\n2 h ago: lost a fight with Sanaki\n\n5 h ago: heard strangers to the east\n\nMonthly summariesthe past year\n\nLast month: groomed most with Mbelo; backed by Rukaso 6 times; made up with Sanaki twice\n\nYearly summarieskept for life\n\nLast year: climbed to rank 3; closest friend Mbelo\n\nWhen a model decides for a chimp, it gets a small slice of this: how the chimp feels, who's around, and a few memory lines about them.\n\nDaily routine\n\nThe simulation keeps real equatorial time: sunrise and sunset just before seven, about twelve hours of daylight all year. The timeline shows what most chimps are doing across a typical day.\n\nLight\n\ndaylight\n\nCalls\n\nDawn chorus06:24–07:24 · everyone joins in with pant-hoots, their loud rising “hoo-hoo-HOO” that carries far\n\nDusk chorus18:00–19:00 · one more round before bed\n\nDoing\n\nAsleepuntil about 06:50 · still in last night's nests\n\nBreakfast, mostly fruit06:50–11:30 · feeding at fruit trees and moving between them\n\nSiesta, grooming11:30–14:30 · resting through the heat, grooming friends\n\nMore feeding14:30–17:50 · back to the trees\n\nBuild a nest, sleepfrom about 17:50 · a fresh nest, then sleep\n\nPatrols\n\nA border patrol may set out08:00–15:30 · every week or two, a group of males heads for the border\n\nStorms\n\nThunderstorms most likely12:30–19:00 · about two-thirds of the rain falls in the afternoon\n\n0608101214161820\n\nRainy seasons\n\nKibale gets rain from March to May and again from September to November, with drier spells in between. Fruit follows the rain about six weeks later.\n\nFigs are different: each tree fruits on its own schedule, so some figs are usually ripe when little else is.\n\n15 °Cat dawn\n\n24 °Cafter lunch\n\nRain per month, mm, as simulatedwet season\n\n62J\n\n78F\n\n150M\n\n185A\n\n160M\n\n78J\n\n60J\n\n110A\n\n155S\n\n185O\n\n180N\n\n95D\n\nAbout 1,650 mm a year. Every new forest opens at 06:30 on 28 September, in the middle of the rains.\n\nNight. The chimps asleep in fresh nests, grouped into nesting parties.\n\nA storm. Most of the party hunches down to wait it out, and Kidoti did a rain display (see the log). Tavuni, still reacting to a stranger's call, patrols.\n\nDecisions\n\nNothing is scripted. Whenever a chimp finishes what it was doing, or something interrupts it (a charge, a storm, a stranger's call), it stops and makes a fresh decision. This repeats for every chimp, all day.\n\nSees and hearsonly what's nearby, no peeking at the map\n\nRemembersfriends, fights, fruit trees, water\n\nLists optionsonly moves that are possible right now\n\nPicks onethe built-in rules or a decision model\n\nActseats, grooms, charges, naps…\n\nWorld changesbellies, bonds, ranks, everyone nearby\n\nEvery chimp, all daythen back to step 1\n\n…and back to step 1.\n\nRules\n\nThe built-in rules score each option with a simple recipe (hungry plus fruit nearby means go eat) and take the top one. They run every chimp the decision model isn't running.\n\nAsync default\n\nThe forest keeps moving while the decision model thinks. If an answer comes back late, the rules fill in and the chimp carries on.\n\nLockstep\n\nThe forest pauses at every model decision and waits for the answer, so every pick lands exactly when it was asked.\n\nOutside the app, test runs hand the same chimps to other deciders: trained adapters, the stand-ins that imitate them in long runs, Jev, and rules that keep their plan.\n\nGLiNER2.5-Decide\n\nIn the app, the decisions the rules don't make come from fastino/GLiNER2.5-Decide, a small classifier that picks from a list of options it's given, running locally. It is one of several deciders ChimpBench tests: the built-in rules, three trained adapters for GLiNER (baseline, aggressive and collaborative), small stand-in networks that imitate those adapters in multi-year runs, and Jev (TypeSafe System One) as an offline test arm. The free part of the decisive test is done; the paid Jev part is approved and pending.\n\nWhat it is\n\nIts model card calls it a specialist classifier for operational decisions: you give it some text and a set of labels at run time, and it scores the labels. It doesn't chat, reason out loud or answer open questions. That's exactly the shape of a chimp's next move: one situation, a short menu, pick one.\n\nWhy small and local\n\nA chimp decides every few minutes of forest time, and dozens of them can be deciding at once. The model has to keep up, run on a laptop with no account, no network and no bill, and only ever choose, never invent. A big chat model would be slower and could make up moves that don't exist.\n\nModel\n\nfastino/GLiNER2.5-Decide\n\nSize\n\n340M parameters\n\nBuilt on\n\nDeBERTa-v3-large encoder\n\nLicense\n\nApache 2.0\n\nRuns on\n\nApple silicon GPU (MPS), fp16\n\nPer decision\n\n~500 input tokens, 0.25–0.4 s\n\nOptions per call\n\n2 to 8\n\nInput limit\n\n1,280 tokens, never truncated\n\nobserve()\n\nA private snapshot of what this chimp sees, hears, feels and remembers. Nothing it couldn't know.\n\nLegal options\n\nUp to 8 moves that are possible right now, each written as a plain sentence, always including the rules' pick.\n\nGLiNER scores\n\nThe model scores every option against the snapshot, in about 0.4 s.\n\nCheck\n\nThe server validates the answer; the simulation re-checks the move is still legal.\n\nAct and trace\n\nThe chimp acts. The trace keeps every score and what the rules would have picked.\n\nLate answers\n\nIn async mode the forest doesn't wait. A late answer is checked against the newer state and applied only if the move is still legal; after six forest minutes the rules decide instead.\n\nLockstep\n\nThe clock stops on the exact tick a model-driven chimp starts waiting and resumes when the answer lands. After 6 real seconds the rules take over, so it can't hang.\n\nTrained adapters and stand-ins\n\nThree LoRA adapters, trained on labelled decisions, change which option GLiNER picks: a field-expert baseline, an aggressive temperament and a collaborative one. They run in headless test runs, not in the app. Stand-ins, small networks fitted to imitate each adapter, run the five-year tests at rules speed.\n\nOffline test arms\n\nJev runs only in offline tests. The free arms (rules, rules that keep their plan, a utility formula and random choice) have run on five fresh worlds. The paid Jev arms are approved with a $10 cap and haven't run yet.\n\nWhat the model actually read418 tokens · 478 ms\n\nmeTavuni, adult male, 27 y, West community; the alpha, rank 1 of 8 males; mood calm; timid; solitary; playful\n\nfeelingmild hunger\n\nnow07:19 dawn; thunderstorm, 16 °C; party of 5 with 2 adult males\n\nnearbySanaki: my ally, adult male, ranks below me, 2 m, close bond · Yerubi: juvenile female, 2 m · 2 more\n\nmemoriesJust now: a heavy storm broke\n\neventsA heavy thunderstorm is pouring down\n\n“You are a field primatologist. Choose what this wild eastern chimpanzee would most plausibly do next, given only what it perceives, feels and remembers…”\n\nSit hunched through the storm AI pickrules pick57%\n\nGroom my ally Sanaki to keep his support17%\n\nGroom Yerubi, juvenile7%\n\nPlay with the young Yerubi6%\n\nSit and rest nearby4%\n\nForage on leaves and pith nearby4%\n\nCharging display to assert my alpha status3%\n\nTravel toward Rukaso's pant-hoots3%\n\nThe Mind tab. Tavuni is at the edge of his range with one companion. GLiNER picks rest (62%) where the rules would have foraged, and the history shows it disagreeing with the rules now and then. Every row keeps the rules' score next to the model's.\n\nWhat the first real runs showed\n\nThe untuned model's first 3,546 logged decisions\n\nIt matches words\n\nIt scores an option largely by how its text relates to the situation's text. So each need and each option is phrased in the same words (\"food\", \"water\"), and only active needs are named: a cue that's always there nudges every decision the same way.\n\nHungry, before7%\n\nHungry, after60%\n\nThirst was its weak spot\n\nThe same wording barely moved thirst: very thirsty chimps almost never picked the stream, even when it was offered. That led to the fix below.\n\nThirsty, before0.3%\n\nThirsty, after7%\n\nStrangers trigger a response\n\nWhen a chimp had just heard strangers, real or from the playback experiment, most picks became a response: flee, call back, patrol toward them, display or charge.\n\nAll quiet9%\n\nStrangers heard69%\n\nIt nests at dusk\n\nBetween about 17:50 and 19:30, with a nest on the menu, it built one more than half the time. At midday a nest is hardly ever even offered.\n\nDusk55%\n\nIt agrees with the rules about half the time\n\n48% across that log, and 40–65% across runs. When they differ it leans social: grooming was its most common pick.\n\nSame pick48%\n\nRepeated facts bias it\n\nRepeating a partner's details made it favor that partner, and \"currently mate-guarding\" next to a \"Mate-guard…\" option locked it into continuing (7 of 36 contexts at 90%+, versus 2 without the echo). Now each fact appears once, and ids never reach the model.\n\nThe thirst fix\n\nTwo changes. The state now names every urgent need, not just the strongest (a very thirsty chimp that was slightly hungrier used to hear only \"hunger\"). And when a bodily need is urgent, the matching option repeats it in the state's own words: \"water, eases severe thirst, needed now\".\n\n120 real contexts from the log were replayed through the model before and after the fix. When thirst is the stronger need, 10 of 12 chimps now drink; the misses are hunger winning. Moderate thirst still rarely drinks, but the rules rarely drink at that level either.\n\n120 replayed real contextsbeforeafter the fix\n\nVery thirsty → drink30% → 50%\n\nHungry → forage75% → 94%\n\nStrangers → respond63% → 75%\n\nDusk → nest50% → 63%\n\nStorm → shelter56% → 56%\n\nThe replayed set has its own before numbers, so they differ from the log above.\n\nSociety\n\nSocial events are not scripted. Rank, alliances, conflict and reconciliation, territories, hunting, family life and parties each come from a small set of rules that run for every chimp, every day. Where a rule follows a field study, the study is cited.\n\nRank and alphas\n\nMales rise by winning contests, scored with Elo ratings (Neumann et al. 2011), and lower-ranked chimps greet higher ones with a pant-grunt, a breathy \"you're the boss\". The alpha usually falls to a rival with friends. Females queue by age and tenure instead (Foerster et al. 2016).\n\nAlliances\n\nChimps who groom each other, back each other up and share meat become allies. When a friend gets charged, allies may jump in, and that's how a smaller male can topple a stronger one.\n\nFights and making up\n\nMost conflicts are charges and bluffs; about 1 in 20 gets physical. Afterwards, rivals often reconcile and friends console the loser (Kutsukake & Castles 2004). A fight that never gets patched up leaves tension that fades over a few weeks, modeled on how chimps' relationships vary in value and compatibility (Fraser, Schino & Aureli 2008).\n\nTerritories and patrols\n\nEvery week or two, a group of males patrols the border in silence (Watts & Mitani 2001). Meeting strangers, they count heads: outnumber them and they charge; outnumbered, they slip quietly home. Most encounters are only heard, not seen (Wilson et al. 2012).\n\nHunting and sharing\n\nWhen several males spot red colobus monkeys, a hunt can start. More hunters, better odds, as at Ngogo (Mitani & Watts 1999). The catch gets shared with friends, allies and family who beg.\n\nBabies and families\n\nA baby rides on mom's belly, then her back, and nurses until 4 or 5. The next one comes about five years after the last (Wallis 1997). Death rates by age follow Ngogo's (Wood et al. 2017), and orphans are sometimes adopted by an older sibling.\n\nFemales change communities\n\nAround age 12, most young females move to a neighboring community and start at the bottom of the female pecking order. Males stay home for life.\n\nFood and weather\n\nA community splits into parties, temporary subgroups that merge and split all day: big ones at a fig feast, small ones when fruit is scarce. Storms send everyone to shelter.\n\nKinship. Every matriline as a tree; dashed lines trace genetic sires.\n\nDominance. Male and female ladders by Elo score, per community. Press T in the app.\n\nField experiments\n\nField experiments test how chimps react when they perceive or experience a specific event, such as a stranger's call played from the edge of their range or a model snake on a trail. In the app, the Experiments panel places the same kinds of events in the simulated forest, and the Mind tab shows a chimp's choice just before and just after.\n\nStranger call playback\n\nA hidden speaker plays a strange male's pant-hoot. Do they go look, call back or sneak away? (Wilson, Hauser & Wrangham 2001)\n\nSnake model\n\nA fake viper on the trail. Who spots it, and do they warn the ones who haven't? (Crockford et al. 2012)\n\nFig feast\n\nA giant fig ripens all at once. Watch everyone show up.\n\nStorm\n\nA heavy downpour, right now. Everyone hunkers down; a few males do a rain display.\n\nDrought\n\nA few days with hardly any fruit. Parties shrink and squabbles over food go up.\n\nRemove the alpha\n\nThe boss vanishes. Watch the struggle to fill the top spot.\n\nMonkeys arrive\n\nA group of red colobus moves in overhead. Hunting chance. (Watts & Mitani 2002)\n\nBefore and after\n\nThe Mind tab puts the chimp's choice right before the experiment next to the one right after.\n\nStranger playback, before and after. In this example run, Tavuni was foraging (40% from the app's decision model, GLiNER2.5-Decide). After a stranger's call about 18 m away, the model picks flee at 100% and the rules agree, while Sanaki and Rukaso set off on patrol. Field playbacks found the same pattern: small parties back off, bigger male parties move in (Wilson, Hauser & Wrangham 2001).\n\nArchitecture\n\nOne rule organizes the code: only the simulation changes the world. The 3D view, the sound and the interface read it; the decision loop asks it for a chimp's view and returns a checked choice.\n\nDecision model: which models are being tested?\n\nThe engine\n\nSimulation\n\nEvery chimp's body, mind and relationships, plus the forest they live in. The only part that changes the world.\n\nClock\n\nMoves the world forward in fixed 15-second steps.\n\nSaves\n\nYour simulations, stored in your browser. Reload and carry on.\n\nThe brains\n\nDecision loop\n\nPicks which chimps ask the decision model, sends it what that chimp sees, and hands a checked pick back to the simulation.\n\nDecision modelReads one chimp's view and returns one pick; it never touches the world directly. Several models are being tested.Which models\n\nWhat you see and hear\n\n3D forest\n\nTerrain, trees, sky, rain and every chimp, animated. Read only.\n\nSound\n\nCalls, birds, rain and thunder, placed around your camera. Read only.\n\nInterface\n\nPanels, family trees, rank ladders, experiments and the Mind tab. It sends your controls (speed, experiments) to the simulation.\n\nDecision models being tested\n\nRules\n\nThe built-in baseline: a simple recipe scores each option.\n\nGLiNER2.5-Decide\n\nThe small local model that runs live in the app.\n\nTrained GLiNER versions\n\nThree temperaments, baseline, aggressive and collaborative, tested outside the app.\n\nJev (TypeSafe)\n\nTested offline. The free test is done; the paid test is pending.\n\nMade with TypeScript, three.js for the 3D, Web Audio for sound, and SQLite running in the browser for saves.\n\nChimp data\n\nThe field data used to recreate a realistic forest and realistic chimps: open datasets from long-term studies at Ngogo, Gombe and Taï set the forest and its fruit seasons, and give the results the simulated chimps are checked against.\n\nThe field data usedNgogoGombeTaï\n\nEight open datasets from three sites, 1978 to 2024. One builds the simulated forest; the other seven test the simulated chimps.\n\nLoading the datasets…\n\nEach dataset is used under its licence (CC0 or CC BY 4.0) and credited wherever its numbers appear. Raw records never leave data/raw/: this guide shows counts, shares, rates and normalized maps only.\n\nEvery target, pass or fail\n\nEach square is one field target. Held-out targets were never used for tuning, and most of the scored ones still fail.\n\nLoading the scorecard…\n\nTargets and pass bands: data/targets.json, from 124 cited studies. Verdicts: the latest year-long field-profile runs of the virtual researcher. A tuned, encoded or compromised target never counts as a pass.\n\nBehaviour by behaviourfield bandeach simulated seedoutside the bandM, F: scored per sex\n\nGrooming time, fruit in the diet and male bonds land in the wild bands on the tuning seeds, less clearly on fresh ones. Being improved: hunting success, patrol rate and home range are well outside their bands.\n\nLoading…\n\nEach row has its own scale, from zero. Bands, sites and sources: data/targets.json; values: the tuning seeds.\n\nFruit calendar\n\nRipe fruit is chimps' main food, and its timing affects party size and travel. At Ngogo, researchers recorded ripe fruit on the same marked trees every month from 1998 to 2017 (Potts et al. 2020). The simulated forest (field profile) uses this record directly: its fruit follows the Ngogo months in order, year by year.\n\nTwenty years of ripe fruit at Ngogo\n\nRipe fruit comes in pulses, and no two years repeat. The simulated forest steps through these years in order, month by month.\n\nLoading the phenology record…\n\nData: Ngogo, Kibale National Park, 1998–2017 (Potts et al. 2020, Biotropica; data Dryad, CC0). The ripe-fruit score is the source's crop-weighted index. Monthly aggregates are derived by ChimpBench; not endorsed by the authors.\n\nRipe fruit by species and month\n\nThe eight Ngogo species in the simulated forest fruit at different times, and the figs carry a little fruit almost all year, which keeps some food available when other species have none.\n\nLoading the species records…\n\nShare of monitored stems with ripe fruit, per species and month (same source, CC0). The ninth species, Ficus sansibarica, isn't on the Ngogo transect and uses the fig average.\n\nMovement\n\nThree open datasets record where wild chimps went: GPS fixes from Ngogo (2011–2023), 30-minute focal records from Taï (2013–2016) and 15-minute focal follows from Gombe (2000–2003). The same measures are computed for real and simulated chimps: home-range area and core area (95% and 50% kernels), how use spreads across the range, 30-minute step length, turning angle and how straight a day's path is. The virtual field observer samples the simulated chimps on the same schedules. Status: the way use spreads within a range matches Ngogo. Being improved: ranges are about three times too small for the group size, and day paths double back too often.\n\nSimilar\n\nLoading the comparison…\n\nDifferent\n\nNot comparable yet\n\nRange area: the simulated communities' ranges are at least 2.7× smaller than the band for wild communities of about 22 members (1.8 km² against 5–16 km²), while the way use spreads inside a range matches. Day paths: net distance divided by distance walked is …, against … at Taï and … at Gombe, and simulated chimps reverse direction more often. Parties: about 3 chimps, against about 6 at Taï.\n\nBeing improved. These charts come from an earlier run. Later changes brought turning and party size closer to Taï, but not straightness (0.23 against 0.50): each change that raised straightness made turning or party size worse. Movement work is paused until the next full test run, and the charts will be redrawn after it.\n\nSimulated ranges, to scalesimulatedwildtop of the size-matched band\n\nAgainst wild communities of the same size, simulated ranges are several times smaller (being improved). Ngogo is shown for scale only: it has about nine times as many chimps, so it is not a like-for-like comparison.\n\nLoading…\n\nSize-matched band: data/targets.json T-RNG-1 (Kanyawara, Budongo Sonso and Waibira; Taï). Taï groups of 11–23 members, yearly 95% kernels (Lemoine et al. 2020, CC BY 4.0). Ngogo community kernel, 2011–2014 (Sandel et al. 2026, CC BY 4.0). Simulated: field profile, 95% kernels, five seeds.\n\nEach chimp · NgogoCommunity · NgogoTerritory · Taï\n\nWhy the maps look abstract. Chimp location data can help poachers, so it's handled carefully and this guide never shows where any chimp actually was. Each map is centred on its range, scaled by its own range radius (r), rotated to its long axis and mirrored. Only the shape is left, and shape is what's being compared.\n\nSix full days: Gombe vs simulatedDifferent\n\nMeasured in range radii, a simulated day covers more than twice the distance of a Gombe day but ends about as far from its start, because it doubles back more often (being improved).\n\nLoading the paths…\n\nSix full days per side at evenly spaced straightness ranks, so neither side is cherry-picked. Real: Gombe National Park, Kasekela community, 2000–2003, 30-minute records from 15-minute focal follows (Pusey & Schroepfer-Walker 2013; data Dryad, CC0). Paths are re-drawn from the records after privacy normalization; not endorsed by the authors.\n\nDistance covered in 30 minutes\n\n…\n\nTaï, western chimpanzeesGombe, eastern chimpanzeesSimulateddotted: median\n\nTurning between 30-minute stepsDifferent\n\nAt both wild sites, most turns between 30-minute steps are small. Simulated chimps make more sharp reversals (being improved).\n\nHow much the ranges of Ngogo's two halves overlap, year by year (1 = the same range, 0 = none). The community polarized in 2015 and was two groups by 2018. No simulated community has split yet, so the simulated lines stay flat.\n\nNgogo, West vs CentralSimulated: best split inside one communitySimulated: two neighbouring communities\n\nHow the day is spentshare of 30-minute records\n\nThe daily budget is close to Taï's: a bit more feeding, a bit less rest.\n\nLoading…\n\nTaï only records rest, travel and feeding, so grooming and other social time count as rest here.\n\nA boundary patrol is a group of males walking the edge of their range, stopping to listen and sometimes crossing into a neighbour’s range. Patrol records from Ngogo, Gombe and Taï give how often patrols happen, who joins and how long they last, and the simulated patrols are measured the same way.\n\nLoading the comparison…\n\nPatrols by month of the yearPer monthShare of the year\n\nGombe is the fair seasonal reference: its patrols come from daily follows all year round, and they happen in every month. Ngogo's monthly counts mostly show when observers were in the field.\n\nLoading the patrol records…\n\nEach site's patrols by calendar month, averaged over its study years. Gombe 1978–2007 (Massaro et al. 2022; data Dryad, CC0). Ngogo 1996–2015 (Langergraber et al. 2017; data Dryad, CC0). Monthly counts only; no dates or names.\n\nPatrols by the numbers\n\nMeasure\n\nNgogo\n\nGombe\n\nBand\n\nSimulated\n\nLoading…\n\nNumbers computed from open records say open data; the rest are published figures. Each band was fixed before any run of the corrected patrol model. Ngogo's community is far bigger than the simulated ones (24–44 males), so shares and rates are the fair comparison, not head counts.\n\nWho joins a patrol\n\nIn a small community most males go on most patrols. In Ngogo's huge one, the typical male joins about a third.\n\nLoading…\n\nOne dot per male, anonymous. Gombe (Massaro et al. 2022, CC0): share of observed patrol opportunities joined over the whole study. Ngogo (Langergraber et al. 2017, CC0): share of the patrols a male was listed on that he joined, males listed on 20 or more.\n\nAdvance or retreat after a border stop\n\n…\n\nTaï, western chimpanzeesSimulatedshaded: the band\n\nWhere border stops happen\n\n…\n\nTaï border stops\n\nBefore and after the Ngogo split\n\nTerritory gained after killings\n\nPatrols vs ordinary visits to the edge\n\nGombe, 1978–2007: patrols against ordinary visits to the edge of the range, seen on the same daily follows.\n\nMeasures not yet comparable\n\nOnly summaries leave the raw files. The patrol records name individual chimps and give exact dates. This guide publishes counts, shares, rates and distances measured in range radii, never a name, a place or a day.\n\nChimpBench can hand the same chimps to different decision policies and score each one with the same virtual field researcher, against the same wild bands. So far, the best result comes from the simplest change: the built-in rules, told to keep their plan and to choose with some randomness.\n\nRules\n\nThe built-in recipe that runs the app. It was tuned on travel share, party size and day range, which gives it a home advantage. Even so, in the free test its chimps travel too much, groom too little and walk too far each day.\n\nUntuned GLiNER2.5-Decide\n\nThe model as downloaded. It leans social: over three observed days its chimps groomed 28% of daylight (wild: 8–18%) and walked 0.15 km a day (wild: 1.5–3.5 km). When aggression was on offer it chose it about five times as often as the field-expert labels (0.29 against 0.06).\n\nTrained adapters\n\nThree small add-ons to GLiNER (LoRA adapters), trained on the same 900 decisions with different labels: a field-expert baseline, an aggressive temperament and a collaborative one. The labels come from an AI labeller following a written rubric, not from field data. On 200 held-out decisions, each adapter matches its own labels far better than the untuned model (0.74, 0.72 and 0.81 against 0.40–0.54), and the temperaments separate clearly.\n\nJev (TypeSafe System One)\n\nA hosted model that turns a described situation and a set of options into probabilities. Driving one community at a time for three days, it chose less aggression and more grooming than rules, and picked social options 64% of the time when they were offered (expert labels: 36%). A paid decisive test is approved but hasn't run.\n\nRules + plan\n\nThe same rules with two changes. A chimp keeps its current plan until something important changes, such as a need crossing a threshold, the time of day moving on, an interruption, the plan ending, or 90 minutes passing. And it picks from the rules' scores with some randomness instead of always taking the top one. In the free test this cut the gap to wild chimps by 55%, on all five worlds.\n\nTwo controls: a formula and random choice\n\nThe utility formula scores each option by how much hunger, thirst, tiredness and loneliness it relieves per hour, walk included, using the simulation's own rates. No model is involved. Random choice picks any legal option. Random choice beat rules on distance, but its chimps went hungry, so it doesn't count.\n\nWhat the model runs showed\n\nReal models: 3-day runs. Five-year runs: stand-ins.\n\nHungry chimps eat leaves instead of walking to fruit\n\nThis is the adapters' main failure. At strong or severe hunger, with leaves the only food nearby and a remembered fruit tree on the menu, they eat leaves in place 97–100% of the time and never make the trip. Rules make it 14–27% of the time. Leaves feed about half as fast as ripe fruit, so a chimp living on leaves has to feed about 75% of the day. Share that walk to the fruit, at strong hunger:\n\nRules14%\n\nAdapters0%\n\nThey feed too much and travel too little\n\nOver three observed days the adapters' chimps fed 59–71% of daylight (wild: 33–50%). Over five simulated years, run with stand-ins that imitate each adapter, the trained communities fed 71–82% of daylight and travelled 2–9% (wild: 12–25%). Adding randomness can't fix this for the adapters: they put nearly all their weight on one option, so it shifts their choices by 2 points or less.\n\nThe temperaments are real\n\nThe aggressive adapter charged about four times as often as rules and was the only one to cause major injuries. The baseline and collaborative adapters made no charges at all, which is probably less aggressive than wild males. The collaborative one groomed about 3.5 times as often as rules. Charges per adult per day:\n\nRules2.1\n\nAggressive8.7\n\nUntuned3.9\n\nBaseline0\n\nCollaborative0\n\nJev0.1\n\nJev and untuned GLiNER barely travel\n\nJev picked a trip in 0.2% of the decisions where one was offered. Picking in proportion to its probabilities, instead of always taking the top one, would still send its chimps travelling in only about 1.5% of decisions. Untuned GLiNER's chimps walked 0.15 km a day.\n\nThree observed days with the real modelsfield map, 2 seeds (rules: 3)\n\nMeasure\n\nWild band\n\nRules\n\nUntuned\n\nBaseline\n\nAggressive\n\nCollaborative\n\nFeedingshare of daylight\n\n33–50%\n\n46%in band\n\n41%in band\n\n71%above\n\n59%above\n\n62%above\n\nTravelshare of daylight\n\n12–25%\n\n26%above\n\n7%below\n\n13%in band\n\n16%in band\n\n9%below\n\nGroomingshare of daylight\n\n8–18%\n\n7%below\n\n28%above\n\n7%below\n\n10%in band\n\n20%above\n\nMale day rangekm walked per day\n\n1.5–3.5 km\n\n3.6 kmabove\n\n0.15 kmbelow\n\n1.2 kmbelow\n\n2.1 kmin band\n\n2.2 kmin band\n\nParty sizechimps together\n\n3–9\n\n2.8below\n\n3.3in band\n\n2.8below\n\n2.05below\n\n4.5in band\n\nMeans over the seeds, after a 180-day rules burn-in. Seed ranges are wide (aggressive grooming ran from 4% to 16%), so read directions, not ranks. These runs used an older copy of the code than the free test below, so compare within a table, not across.\n\nFree test: distance from the wild bandslower is better\n\nKeeping the plan and adding some randomness cut the gap to wild chimps by 55%, and beat rules on all five worlds. The paid Jev arms haven't run yet.\n\nRules2.04\n\nUtility formula1.83\n\nRandom choice1.37\n\nRules + plan0.93\n\nJev as shipped: pendingJev with facts and randomness: pendingSame, facts shuffled: pending\n\nThe distance sums nine rows: feeding, travel and grooming for each sex, rest, party size and male day range. A row counts 0 inside the wild band; otherwise it counts the gap to the nearest edge, divided by the band width. Mean of five fresh worlds (seeds 6501–6905), all three communities on one policy, five scored days after a 180-day rules burn-in, measured on simulation truth.\n\nPre-registered before any run, after a four-judge review of two Jev designs: docs/staging/jev-decisive-test.md. Results: artifacts/decide-ft/jev-test/free-arms.md. Lactating females are left out of the score until a known energy problem in the simulation is fixed (they go hungry even under rules), and are reported separately.\n\nThe free test, policy by policy5 worlds each\n\nPolicy\n\nDistance\n\nBetter than rules\n\nMedian hunger\n\nResult\n\nRulesas shipped\n\n2.04\n\n—\n\n0.64\n\nControl\n\nRules + plankeeps its plan, picks with some randomness\n\n0.93\n\n5 of 5 worlds\n\n0.66\n\nBest so far\n\nUtility formulaneeds relieved per hour, no model\n\n1.83\n\n3 of 5\n\n0.39\n\nSmall gain\n\nRandom choiceany legal option\n\n1.37\n\n5 of 5\n\n0.86\n\nNon-viable\n\nJev as shippedcurrent packet, top pick\n\npending\n\npending\n\npending\n\nPending\n\nJev with facts and randomnessfood values and walking times in the options\n\npending\n\npending\n\npending\n\nPending\n\nSame, facts shuffled1 world: does Jev read the facts?\n\npending\n\npending\n\npending\n\nPending\n\nHunger runs from 0 (full) to 1, for all adults. A policy is non-viable if its median adult hunger is more than 0.10 above rules', or its lactating females' median reaches 0.95; random choice fails both (0.86 and 0.97). Rules miss on travel, grooming and male day range. Rules + plan brings travel and day range into the band, but grooming stays low and parties stay a little small.\n\nNo training on Jev's answers. TypeSafe's customer agreement (§2.3(b)) forbids using Jev's outputs to train a model that imitates it. So Jev can't be distilled into a fast stand-in for long runs without TypeSafe's written permission, and its answers carry a do-not-train marker.\n\nStill pending\n\nThe paid Jev test. Approved with a hard $10 cap, not run yet: it waits for a spending guard tested against a fake server. Jev counts as helping only if its version with facts and randomness beats rules by at least 0.10 on average and on 4 of 5 worlds, beats both rules + plan and the utility formula by at least 0.05 on 4 of 5 worlds, and keeps its chimps fed.\n\nRules + plan as the default policy. Approved and pre-registered as stage C13. It can be switched off, and the combined proof run will judge it. That run hasn't happened yet.\n\nLong runs of the trained adapters. The five-year numbers come from stand-ins: small networks fitted to imitate each adapter, which agree with it on 67–77% of decisions in worlds they never saw. They are not the adapters themselves.\n\nCredits\n\nThe tools, decision models, sound and field data behind ChimpBench, and the studies it draws on, each with its licence, are listed on the Credits page.", "url": "https://wpnews.pro/news/chimpbench-chimpanzee-population-simulator", "canonical_source": "https://chimpbench.pages.dev/about", "published_at": "2026-09-30 23:42:07+00:00", "updated_at": "2026-09-30 23:48:37.137857+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-agents", "ai-research"], "entities": ["ChimpBench", "Kibale National Park", "Ngogo", "Kanyawara", "Mitani", "Watts", "Amsler", "Koruza"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/chimpbench-chimpanzee-population-simulator", "markdown": "https://wpnews.pro/news/chimpbench-chimpanzee-population-simulator.md", "text": "https://wpnews.pro/news/chimpbench-chimpanzee-population-simulator.txt", "jsonld": "https://wpnews.pro/news/chimpbench-chimpanzee-population-simulator.jsonld"}}