cd /news/ai-agents/my-mergetober-story-building-a-bambo… · home › topics › ai-agents › article
[ARTICLE · art-147202] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

"My Mergetober story: building a BambooHR connector for cognee"

A developer built a BambooHR connector for cognee, the open-source knowledge-graph memory layer for AI agents, after proposing a change-feed-based merge strategy instead of the full-replace approach used by the existing Notion connector. The connector syncs employee directories and company files, using BambooHR's change feed with a success-only cursor to avoid data loss on crashes, and was tested locally against a free BambooHR trial with Ollama running llama3.1:8b and nomic-embed-text.

by read5 min views3 publishedOct 7, 2026

This Mergetober I wanted to do more than fix a typo. I wanted to build something real, get it reviewed, and actually understand every line of it. I ended up building a BambooHR connector for cognee, and this post is the story of how it went, mistakes included.

cognee is an open source project that turns your data into a knowledge graph AI agents can search, basically long-term memory for agents. It uses "connectors" to pull data in from tools like Notion, Gmail and Google Drive.

While browsing the issues I found #4795: add a connector for BambooHR, an HR platform. The idea was simple: if your company's employee directory and handbooks are in BambooHR, your AI agent should be able to answer "who works in Sales?" or "how many days of leave do we get?"

The issue suggested copying the existing Notion connector. Before commenting, I read through BambooHR's API docs and noticed something important: BambooHR has a change feed, an endpoint that tells you exactly which employees were added, updated or deleted since a given time.

The Notion connector re-downloads everything every run and replaces the old data, because Notion has no such feed. If I copied that approach but only fetched changes, every employee who didn't change would get wiped from memory. So in my comment I proposed using the change feed with a merge strategy instead, like the Gmail and Google Drive connectors do. I also asked what should happen to employees who leave the company.

That comment got me assigned. 🎉

My first hour was mostly Git mistakes:

git init in the parent folder instead of inside the cloned repo. Not my proudest moment, but it's the kind of thing you only do once.

I didn't have access to a real company's BambooHR (and wouldn't want to test on one anyway), so I signed up for a free BambooHR trial. It comes preloaded with around 118 fake employees and about 20 company files like handbooks and policies, which turned out to be perfect for testing.

I spent a while just calling the API by hand to see what it actually returns. That's where I learned two things that shaped the whole design:

1970-01-01, every employee comes back as "Inserted". One code path handles both the first sync and every sync after it. The connector syncs two things: the employee directory and company files (PDF, TXT, MD, CSV).

For employees, it uses the change feed and stores BambooHR's own timestamp as a cursor for the next run. That cursor is only saved once the whole run succeeds, so a crash halfway doesn't skip anything. Deleted employees are removed, and terminated employees (which BambooHR marks as Inactive) are forgotten by default too.

Files were trickier because BambooHR has no change feed for them. So each run lists all files, compares them with the previous run to spot deletions, and lets cognee skip files whose content hasn't changed.

I didn't want to pay for API calls while experimenting, so I ran everything locally with Ollama, using llama3.1:8b as the LLM and nomic-embed-text for embeddings. This is where most of my debugging time went.

Ollama models "downloaded" but didn't exist. ollama pull said success, yet ollama list was empty. It turned out Ollama 0.40.0 stores model manifests as symlinks, and Windows 11 refused to follow them. Replacing the two symlinks with real copies of the files fixed it.

Windows' 260-character path limit. cognee's vector store creates deeply nested folders, and my long project path pushed it over the limit. Moving the data to a short folder fixed it.

The model returned the schema instead of the data. cognee asks the LLM for JSON in a fixed shape, and the small 8B model sometimes replied with the description of the shape instead. Adding this to .env solved it:

LLM_INSTRUCTOR_MODE="json_schema_mode"

Once it all worked, I loaded a few mock employees and a leave policy and asked:

Then I deleted that Sales employee upstream, synced again, and watched cognee's document count drop from 4 to 3. Seeing it forget someone was honestly the most satisfying moment of the project.

Next I ran the connector against my BambooHR trial four times, making changes in between:

Run Employees fetched What happened Stored
1 (first sync) 108 Deleted and inactive sample employees skipped 88
2 20 Picked up my edits 88
3 5 One deleted, one terminated 87
4 2 Caught an employee I terminated while run 3 was still running 86

Run 4 was a nice confirmation that saving the cursor only at the end really works: nothing slipped through.

The trial also surprised me. One of the sample files was a Benefits CSV with per-employee data in it. My field allowlist protects the employee directory, but it can't protect against a file like that. So I added a file_categories option to let you choose exactly which file folders get synced. Another sample PDF turned out to be genuinely corrupt, so I made sure broken files are skipped with a warning instead of crashing the run.

Alongside the live tests I wrote 33 unit tests with a mocked API, so reviewers can run them without a BambooHR account.

I opened the PR… and CI failed immediately. The repo requires a DCO sign-off on every commit, which I'd never heard of. The fix:

git rebase --signoff upstream/main
git push --force-with-lease

If your team uses BambooHR, here's all it takes:

cd packages/connector/bamboohr && uv sync
export BAMBOOHR_COMPANY_DOMAIN=acme   # from https://acme.bamboohr.com
export BAMBOOHR_API_KEY=...           # BambooHR → your name → API Keys
python
import asyncio
import cognee
from cognee_community_connector_bamboohr import bamboohr_source

async def main():
    await cognee.remember(
        bamboohr_source(),
        dataset_name="bamboohr",
        primary_key="id",
        write_disposition="merge",  # important: cognee's default is replace
    )

    answer = await cognee.search(
        query_text="Who works in the Sales department?",
        query_type=cognee.SearchType.GRAPH_COMPLETION,
        datasets=["bamboohr"],
    )
    print(answer)

asyncio.run(main())

Run remember again anytime and only the changes are synced. You can also pass include_inactive=True to keep former employees, or file_categories=["Company Files"] to limit which files are synced.

Thanks to the cognee maintainers for the guidance on the issue, and to the WeMakeDevs community for Mergetober pushing me to ship something real.

Links

── more in #ai-agents 4 stories · sorted by recency
── more on @cognee 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/my-mergetober-story-…] indexed:0 read:5min 2026-10-07 · —