Building a Photo-Based Card Grader: Computer Vision Where the LLM Doesn't Pick the Score A developer behind CardGrade, a phone-photo trading-card pre-grading app, described an architecture in which a border-detection model warps each card side to a flat canvas and slices it into 16 inspection zones, with a vision model only describing defects while a deterministic rubric assigns PSA, BGS and CGC estimates. The pipeline returns results with confidence scores in about 60 seconds, and a V2 cropper abandons the earlier 4000px upscale because interpolation acted as a low-pass filter that erased fine surface scratches while preserving low-frequency corner and edge geometry. Trading-card grading looks like an image-classification problem until you try to build it. A grade is driven by several physical properties at once: how centered the art is, how sharp the corners are, whether an edge is whitening, whether the surface has a fold. Each one lives at a different scale in a phone photo. This is a write-up of how we structured the pipeline behind CardGrade https://cardgrade.io , an app that pre-grades cards from phone photos. Its engine, CGI Vision AI, returns estimated PSA, BGS and CGC grades with confidence scores in about 60 seconds. It's an architecture post, not a benchmark post. I'm not quoting an accuracy number, because a photo-based result is an estimate, and I'd rather say so than dress it up. react-native-vision-camera for capture The grading request is small. The web app stores the full-resolution originals and sends the CV service a payload with a grading ID, front and back image URLs, and a webhook URL. The service replies 202 , does the work, and calls the webhook with structured results. The reasoning: the CV service is the only place where the original pixels, the detected card border, and every consumer of the crops meet in a single request. If the web app also cropped images, we'd maintain two implementations of the same geometry, and they would drift. The first model is a border detector. It finds the card's outer edge and the inner edge of the printed border. Those two quadrilaterals drive everything downstream. From them we compute a perspective transform and warp each side onto a flat canvas. Small details bite here: x + y , top-right maximizes x − y , bottom-right maximizes Once the card is flat and axis-aligned, zones become slices of one canvas per side: a fixed fraction of the card width for each corner, a thin band for each edge, and the inner art area for surface. Four corners and four edges on two sides gives the 16 inspection zones CGI Vision AI reports on. Centering and surface are assessed alongside them. An early version of the pipeline upscaled small uploads to roughly 4000px on the long edge before doing anything else. The border model was tuned around that scale, and many absolute-pixel thresholds were calibrated to it. The catch is that upscaling is interpolation, and interpolation is a low-pass filter. Fine scratches are high-frequency detail, and they're gone before a crop is cut. Corner and edge geometry is low-frequency, so it survived. Surface defects didn't. Removing the upscale outright would change three things at once: the border model's input distribution, the polygon coordinate space, and every absolute-pixel constant. So the V2 cropper uses a two-source model: The V2 cropper is built around this. It is a Python port of the crop geometry with a versioned geometry registry, checked against the original TypeScript implementation on real-photo fixtures to sub-pixel parity. Zone crops are written under a -v2 naming scheme with a geometry sidecar, so every crop can be traced to the exact geometry that produced it. Anything measurable gets measured with OpenCV instead of guessed: For defects that are hard to hand-engineer, a vision model examines each zone crop and describes what it sees, with a severity of none, minor, moderate or severe. The division of labor is strict: the AI never picks scores, it only describes. A deterministic rubric converts those observations and the CV metrics into numbers. That split means you can trace why a card scored what it did, and changing a scoring rule doesn't mean re-prompting anything. Caps are explicit too: a creasing finding caps the surface score by severity. Each subgrade uses the weakest zone : corners is the minimum of the four corners, edges the minimum of the four edges, with no averaging. The overall grade is a weighted blend of the four subgrades, rounded to the nearest half point. Centering carries the least weight and surface the most. If centering can't be measured, it is dropped and the remaining weight is redistributed rather than filling in a made-up number. On top of the blend sit weakest-link caps, so one badly damaged pillar can't be averaged away. Those weights are our model of the process, not anything PSA, BGS or CGC publish. Results are estimates, not official grades, and CardGrade is independent of all three. Surface is where a single phone photo is weakest. Some professional systems build a surface topology map from multiple lighting angles. A phone photo gives you one lighting condition. Corners and edges have measurement channels. Creases are harder, because a fold is a three-dimensional event and a photo is flat. A model's opinion alone can go wrong in both directions: a real fold can be missed, and a print line can be read as a crease. So crease handling is layered: The member-facing product is native iOS and Android on React Native and Expo. JS-only changes ship over the air with EAS Update. The runtime version is tied to the app version, so an OTA only reaches matching binaries. When a change alters the client/server contract, keep the server backward compatible with clients that may never update. If you want to see the result, it's at cardgrade.io https://cardgrade.io . Questions about the pipeline are welcome in the comments.