{"slug": "i-built-a-self-hosted-ai-code-reviewer-for-github-prs-with-spring-boot", "title": "I Built a Self-Hosted AI Code Reviewer for GitHub PRs with Spring Boot", "summary": "A developer built CodeGuard AI, a self-hosted AI code-review application using Java 17+ and Spring Boot that inspects GitHub pull requests, analyzes them for bugs, security, performance and best-practice issues, and posts categorized findings back to GitHub. The tool supports multiple AI providers — OpenAI, Google Gemini and Ollama — so the AI layer is replaceable, and it returns structured output including an overall score, AI confidence, severity-categorized findings, code locations and suggested fixes. In testing, it flagged a GitHub token override configuration issue and a repository-name case-sensitivity problem.", "body_md": "Why I built this\n\nCode review is one of the most important parts of a software development workflow, but reviewing every Pull Request manually can become repetitive.\n\nI wanted to experiment with a different approach:\n\nWhat if an AI reviewer could inspect a GitHub Pull Request, identify potential problems, explain them, and post the review back to GitHub?\n\nThat led me to build CodeGuard AI.\n\nIt is a self-hosted AI code-review application built with Java 17+ and Spring Boot.\n\nThe goal wasn't to replace human code review.\n\nInstead, the idea was to create an additional automated layer that can identify potential issues before a human reviewer spends time going through the PR.\n\nWhat CodeGuard does\n\nThe workflow is essentially:\n\nGitHub Pull Request\n\n        ↓\n\nGitHub API / Webhook\n\n        ↓\n\nSpring Boot Review Service\n\n        ↓\n\nFetch PR information + changes\n\n        ↓\n\nAI Provider\n\n        ↓\n\nBug / Security / Performance / Best Practice Analysis\n\n        ↓\n\nQuality + Risk + Confidence\n\n        ↓\n\nGitHub Comments\n\n        ↓\n\nReview History & Trends\n\nThe application supports multiple AI providers:\n\nOpenAI\n\nGoogle Gemini\n\nOllama\n\nThat makes the AI layer replaceable instead of tying the application to a single provider.\n\nThe application dashboard\n\nThe main dashboard provides two ways to start a review.\n\nA review can be triggered manually by providing a repository and Pull Request number, or the application can work with the GitHub workflow.\n\nIt also exposes the major review categories directly in the interface:\n\nSecurity Review\n\nPerformance\n\nBest Practices\n\nBug Detection\n\nThe manual-review interface was particularly useful during development because it allowed me to test the complete review pipeline without depending on a live webhook every time.\n\nStarting a review\n\nFor example, I can provide a repository and PR number and start the review.\n\nThe application then moves through the review pipeline.\n\nThe UI exposes the progress of the operation instead of making the user wonder whether the review is still running.\n\nThe basic stages are:\n\nReading PR\n\n    ↓\n\nFinding bugs\n\n    ↓\n\nChecking security\n\n    ↓\n\nGenerating review\n\nThis also highlights an important part of AI integrations that can easily be overlooked:\n\nAI operations are asynchronous from the user's perspective.\n\nA review may involve API calls, fetching repository data, preparing the PR changes, sending the analysis request, processing the response, and finally publishing the result.\n\nWhat does the AI review actually produce?\n\nAfter the analysis is complete, CodeGuard generates a structured review rather than returning one large block of AI text.\n\nFor one of my test Pull Requests, the result included:\n\nOverall score: 8/10\n\nAI confidence: 90%\n\nCritical issues: 0\n\nSuggestions: 2\n\nMerge recommendation\n\nCategorized findings\n\nCode locations\n\nSuggested fixes\n\nThis makes the output much easier to consume than simply displaying an LLM response.\n\nThe reviewer can immediately see the overall result and then drill down into individual findings.\n\nFindings are categorized\n\nEach finding is classified so that developers can focus on the type of problem they care about.\n\nFor example:\n\nCritical\n\nHigh\n\nMedium\n\nLow\n\nSecurity\n\nPerformance\n\nBest Practice\n\nIn one test, CodeGuard identified a potential configuration issue involving GitHub token overrides.\n\nIt also detected a repository-name case-sensitivity issue.\n\nThe important part here isn't just detecting an issue.\n\nThe system also provides:\n\nLocation → Explanation → Suggested fix\n\nThat makes the result more actionable for a developer.\n\nPosting the result back to GitHub\n\nOne of the parts I wanted to experiment with was closing the loop.\n\nThe reviewer shouldn't have to stay inside a separate dashboard forever.\n\nCodeGuard can post the review back to GitHub.\n\nThe dashboard therefore shows whether the review has been posted successfully.\n\nThis creates a workflow closer to:\n\nDeveloper opens PR\n\n        ↓\n\nCodeGuard analyzes it\n\n        ↓\n\nFindings generated\n\n        ↓\n\nReview posted to GitHub\n\n        ↓\n\nDeveloper fixes issues\n\n        ↓\n\nHuman reviewer continues the normal review\n\nReview history\n\nAnother feature I wanted was historical visibility.\n\nInstead of treating every PR review as an isolated event, CodeGuard stores review results and exposes them through the review history.\n\nThis allows the application to show:\n\nPrevious Pull Requests\n\nQuality scores\n\nIssues found\n\nRepository information\n\nReview status\n\nReview details\n\nReview trends\n\nThe history also provides a simple trend view.\n\nThe dashboard tracks:\n\nQuality Score vs Issues Found\n\nover previous reviews.\n\nThis is useful because a single review score doesn't tell the whole story.\n\nFor example, if a project repeatedly receives lower-quality reviews or increasing numbers of findings, that pattern may be more interesting than any individual PR.\n\nMaking the AI provider configurable\n\nAnother engineering decision was avoiding a hard dependency on one AI provider.\n\nCodeGuard supports:\n\nOpenAI\n\nGemini\n\nOllama\n\nThe provider can be selected/configured through the application settings.\n\nThis also makes the project interesting for developers experimenting with local models.\n\nFor example, a developer can use a hosted provider during development and experiment with a local Ollama setup when self-hosting is preferred.\n\nCustom review instructions\n\nThe settings also allow custom review instructions.\n\nFor example, a team could specify additional standards such as:\n\nFlag public methods without Javadoc.\n\nFlag endpoints that don't validate input.\n\nFlag direct System.out usage instead of logging.\n\nFlag newly thrown generic exceptions.\n\nThese instructions are added to the review process alongside the built-in review criteria.\n\nThis makes the reviewer more adaptable to a project's own engineering standards rather than relying only on generic AI code-review prompts.\n\nThe Spring Boot architecture\n\nThe backend is built around Spring Boot.\n\nThe main pieces include:\n\nGitHub Integration\n\n        ↓\n\nPR / Diff Retrieval\n\n        ↓\n\nReview Service\n\n        ↓\n\nAI Provider Layer\n\n        ↓\n\nStructured Review Result\n\n        ↓\n\nFinding Classification\n\n        ↓\n\nGitHub Comment / Dashboard\n\n        ↓\n\nReview History\n\nThe project uses:\n\nJava 17+\n\nSpring Boot\n\nSpring Security\n\nGitHub REST API\n\nAI provider integrations\n\nREST APIs\n\nMaven\n\nHTML/CSS/Vanilla JavaScript\n\nThe important architectural decision was keeping the AI provider layer replaceable.\n\nThat means the application isn't fundamentally tied to one model provider.\n\nWhat I learned building it\n\nSending code to an LLM is relatively easy.\n\nBuilding a useful code-review product around that response is considerably more involved.\n\nYou need to think about:\n\nWhat constitutes a finding?\n\nHow should severity be represented?\n\nHow do you identify the relevant code location?\n\nHow do you avoid overwhelming the developer?\n\nHow should findings be presented?\n\nHow should the result get back into the existing GitHub workflow?\n\nAn AI system can be highly confident about a low-severity issue.\n\nSo I wanted CodeGuard to keep concepts such as:\n\nseverity\n\nand\n\nAI confidence\n\nseparate.\n\nThat gives developers more information when deciding which findings deserve attention.\n\nThe purpose of CodeGuard isn't:\n\n\"AI reviewed the PR, therefore merge it.\"\n\nThe more useful model is:\n\nAI performs an additional automated review layer, while humans remain responsible for the final engineering decision.\n\nThat distinction became important as I designed the workflow.\n\nWhat could be improved next?\n\nThere are several directions I'd like to explore:\n\nBetter inline GitHub comments\n\nMore precise code-location mapping\n\nImproved false-positive handling\n\nRepository-level review policies\n\nCustom severity rules\n\nMore granular review configuration\n\nBetter webhook lifecycle handling\n\nAdditional local-model support\n\nMore detailed review analytics\n\nThe project is intentionally designed as a self-hosted foundation that developers can customize, rather than a hosted SaaS product.\n\nFinal thoughts\n\nBuilding CodeGuard taught me that an AI code reviewer is not simply:\n\nGitHub API + LLM = code reviewer.\n\nThe interesting engineering work happens around the model:\n\nGitHub integration → PR processing → review orchestration → structured findings → severity → confidence → developer feedback → history.\n\nThat's where the application becomes an actual workflow rather than just an AI prompt.\n\nWant to explore the implementation?\n\nI built CodeGuard AI as a self-hosted Spring Boot source-code project for developers who want to study, customize, and build on this type of AI code-review workflow.\n\nGet the CodeGuard AI source code on Gumroad: [https://javacoder716.gumroad.com/l/codeguard-ai](https://javacoder716.gumroad.com/l/codeguard-ai)", "url": "https://wpnews.pro/news/i-built-a-self-hosted-ai-code-reviewer-for-github-prs-with-spring-boot", "canonical_source": "https://dev.to/sweety717/i-built-a-self-hosted-ai-code-reviewer-for-github-prs-with-spring-boot-1ocf", "published_at": "2026-09-30 01:29:54+00:00", "updated_at": "2026-09-30 01:46:45.858716+00:00", "lang": "en", "topics": ["ai-tools", "developer-tools", "ai-agents", "large-language-models", "ai-products"], "entities": ["CodeGuard AI", "GitHub", "Spring Boot", "Java", "OpenAI", "Google Gemini", "Ollama"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-built-a-self-hosted-ai-code-reviewer-for-github-prs-with-spring-boot", "markdown": "https://wpnews.pro/news/i-built-a-self-hosted-ai-code-reviewer-for-github-prs-with-spring-boot.md", "text": "https://wpnews.pro/news/i-built-a-self-hosted-ai-code-reviewer-for-github-prs-with-spring-boot.txt", "jsonld": "https://wpnews.pro/news/i-built-a-self-hosted-ai-code-reviewer-for-github-prs-with-spring-boot.jsonld"}}