# I built a PR reviewer that survived its own model being deprecated

> Source: <https://dev.to/danhpaiva/i-built-a-pr-reviewer-that-survived-its-own-model-being-deprecated-2p1k>
> Published: 2026-10-02 16:59:59+00:00

*This is a submission for the [Hacktoberfest Weekend Challenge: Build for a Friend](https://dev.to/challenges/hacktoberfest-weekend-2026-10-01)*

**groq-pr-reviewer-net** — a .NET 10 CLI that reads your `git diff` and returns a structured code review in your terminal, before you ever open the pull request.

```
dotnet run -- --staged
```

That's it. It reviews bugs and correctness, security, performance, and readability.

**The friend I built it for is the developer who codes alone.** No teammate to tag for a second opinion, no budget for a paid review bot — the solo maintainer, the student, the person shipping a side project at 1am. I've been that person on every side project I've ever started.

The problem it solves is specific: the moment right before you commit, when you know you should have someone look at this, and there is nobody to ask. Paid AI review tools answer that with a seat license and a company card. This answers it with a terminal and a free API key.

A CLI has no deployed link, so here is the whole thing end to end.

**1. Setup you can verify without leaking anything**

`--check` reports where the key was loaded from, how long it is, and whether it has the right shape — **without ever printing the key**. One command tells you if you are ready, before you spend a single request.

**2. The open-weight catalogue**

Every model my key can actually reach. Pay attention to what is *missing* from that list — it matters in a moment.

**3. The model I built the whole thing on, dead**

No Llama chat model is left in the catalogue. Notice that the failure **explains itself**: it names the likely cause, links the live model list, and tells you which flag fixes it. That was the one thing I most wanted to get right, because I'd just lost an hour to an error message that told me nothing.

**4. Same diff, one flag different, working review**

Identical command minus one argument. That gap between screenshot 3 and screenshot 4 is the entire argument of this post, and it cost me about forty seconds.

Screenshot 4 is the tool reviewing the commit that implements it. Earlier passes, against earlier versions of this code, caught three things worth showing — all since fixed, which is why you won't find them in the screenshot above.

It found a real bug I had introduced minutes earlier — I'd changed the default model constant and the README, but forgot the `--help` text:

**Inconsistent default model description** – `PrintUsage()` documents the default model as `llama-3.3-70b-versatile`, while the code constant `DefaultModel` is `openai/gpt-oss-120b`. Users may be confused about which model is actually used.

It caught an unguarded JSON access that would throw if the API schema ever shifted:

**Potential JSON-parsing failure in `FetchModelIds`** – the method assumes the response contains a top-level `"data"` array. If Groq changes the schema or returns an error payload without that property, `GetProperty("data")` will throw an unhandled `KeyNotFoundException`.

And it flagged something missing from my own threat model entirely:

**Potential secret leakage** – the diff itself is sent to Groq unchanged; if the diff contains passwords, tokens, or other secrets they will be transmitted to an external service.

That last one became a warning section in the README. **The tool wrote its own security disclaimer.**

And it keeps earning its keep: the run in screenshot 4 flagged that I never dispose the `HttpResponseMessage` in either API call — a socket leak I had walked straight past.

It is not magic. Across two earlier review passes it insisted, both times, that `net10.0` was not a valid target framework and that I was missing a `using System.Linq;`. Both wrong — the project builds clean with zero warnings. The model's knowledge cutoff predates .NET 10, and `ImplicitUsings` already covers LINQ.

Screenshot 4 has its own share: it claims the diff is "printed to stdout before being sent," which simply never happens. Roughly one in four findings is noise.

I put that in the README verbatim, because a tool that oversells itself is worse than no tool:

Treat the output as a fast second opinion, not as truth. Read it the way you would read a well-meaning junior reviewer.

A junior reviewer who is occasionally confidently wrong but catches things you missed is still worth having. Pretending otherwise would be the actual failure.

A C# (.NET 10) CLI that reviews your `git diff` using an open-weight model
(`openai/gpt-oss-120b`, Apache 2.0) running on Groq's ultra-fast inference.

Point it at any git repository and it prints a structured code review in your terminal — before you open the pull request.

```
dotnet run -- --staged
```

*Above: the tool reviewing its own source code.*

`--list-models` plus `--model` gets you moving again without
touching the code.
The whole thing is one `Program.cs` using top-level statements — deliberately small enough to read in one sitting and fork.

The core is unremarkable on purpose, which is the point: an open-weight model behind an OpenAI-compatible endpoint means no SDK, no framework, no abstraction layer to learn.

``` js
const string GroqApiUrl = "https://api.groq.com/openai/v1/chat/completions";

// Open-weight (Apache 2.0) and currently the strongest chat model on Groq.
// Run --list-models if this one is ever retired.
const string DefaultModel = "openai/gpt-oss-120b";
```

The system prompt is plain text, not a framework construct — which is exactly why swapping models costs nothing:

``` js
const string systemPrompt = """
    You are a senior code reviewer. Analyse the pull request diff below and reply in English,
    as bullet points organised into these sections:
    - Bugs and correctness
    - Security
    - Performance
    - Best practices / readability

    Be concise and specific. If a section has nothing worth raising, write "Nothing to flag".
    Ignore trivial formatting changes.
    """;
```

MIT licensed.

**The open-source AI:** [`openai/gpt-oss-120b`](https://huggingface.co/openai/gpt-oss-120b), an **open-weight model released under Apache 2.0**, served on [Groq](https://groq.com) for inference. No proprietary model touches this project.

**The stack:** .NET 10, a single `Program.cs`, and `HttpClient`. No agent framework, no orchestration library, no vector database. The entire dependency list is the .NET base class library.

That minimalism is a design choice tied directly to the open-weight decision. Because open models are served behind the same OpenAI-compatible `/chat/completions` contract, I didn't need an abstraction layer to stay portable — **the HTTP contract *is* the abstraction layer.** A 30-line `HttpClient` call is already provider-agnostic. Anything heavier would have added lock-in rather than removing it, and would have been one more thing for a forker to learn.

The flow is four steps: shell out to `git diff`, truncate at 60k characters to respect the context window, POST it with the system prompt, print the response.

Two things took longer than the happy path:

**Reading streams without deadlocking.** Draining `git`'s stdout and stderr sequentially will hang the process if stderr fills its pipe buffer while you're still reading stdout. Both reads have to be in flight before you wait:

``` js
var outputTask = process.StandardOutput.ReadToEndAsync();
var errorTask = process.StandardError.ReadToEndAsync();
await Task.WhenAll(outputTask, errorTask);
await process.WaitForExitAsync();
```

**Making the setup failure legible.** Which brings me to the hour I lost.

[Groq](https://groq.com) is an inference provider for open models; keys start with `gsk_`. [xAI's Grok](https://x.ai) is an unrelated company with proprietary models; keys start with `xai-`. The names differ by one letter.

I generated a key on the wrong console and spent an hour staring at `401 Invalid API Key`.

So I built the diagnostic I wished I'd had:

``` bash
$ dotnet run -- --check
Key source  : /path/to/groq-pr-reviewer-net/.env
Length      : 56 characters
gsk_ prefix : ok
Model       : openai/gpt-oss-120b
```

It reports where the key was loaded from, its length, and whether it has the right shape — **without ever printing the key itself.** The 401 message now names the `gsk_` vs `xai-` mix-up outright, and the README carries the warning up front.

Small thing. But "build for a friend" means thinking about the person hitting your error message at midnight, and that person has no idea the two companies exist.

I was going to write the usual paragraph about open weights and avoiding vendor lock-in. Then the argument proved itself while I was still building.

**Halfway through, my model was deprecated.**

I'd built the whole thing on Llama 3.3 70B. When I ran the first real end-to-end test against the API:

```
Groq API error (404 NotFound): The model `llama-3.3-70b-versatile`
does not exist or you do not have access to it.
```

Not a single Llama chat model was left in the catalogue.

Here's what a closed API would have cost me: a new SDK, a new request shape, a new auth flow, a rewritten prompt, and a wait on somebody's migration guide. Possibly a dead weekend project.

**What it actually cost me was a flag.**

```
dotnet run -- --list-models      # what can I actually reach right now?
dotnet run -- --model qwen/qwen3.8-27b
```

I moved to `openai/gpt-oss-120b`, re-ran the review, and shipped. Same endpoint, same request body, same prompt. I added `--list-models` during the recovery so the next person hitting a dead model can see their options in one command instead of digging through changelogs.

That is what open innovation made possible here, concretely: **the difference between a flag change and a rewrite.** Interchangeable models behind a common contract mean no single vendor's roadmap can end your project.

The rest of the case holds too, and matters specifically for the friend I built this for:

For someone with no team and no budget, "I can swap the engine myself" isn't an abstract principle. It's whether the tool still works next year.

*Built in a weekend with .NET 10 and an open-weight model that costs nothing to run.*
