5 Best Practices for Building Robust Python AI Libraries A guide on building robust Python AI libraries outlines five practices, including full type coverage with a `py.typed` marker, a `pyproject.toml`-first structure conforming to PEP 621, input and output validation, dependency isolation, and CI enforcement, citing OpenAI's Python SDK and Instructor as reference implementations. The article argues that AI libraries face distinct failure modes — non-schema-guaranteed model outputs, gigabyte-scale dependencies, and unreliable third-party APIs — that standard Python packaging advice does not address. It recommends running both pyright and mypy in CI, as OpenAI's SDK does, to catch differences between type checkers. 5 Best Practices for Building Robust Python AI Libraries This article covers building robust Python AI libraries specifically, Python AI SDK best practices, and what separates a production-ready AI package from one that only survives in its own demo. A well-loved open-source AI wrapper works perfectly in the maintainer's demo notebook. A week after someone else adopts it, three things break in three different ways. A missing API key crashes with a bare KeyError instead of a message anyone could act on. Installing the package for a simple text-classification feature silently pulls in a four-gigabyte deep learning framework nobody asked for. And the test suite only passes when the model provider happens to be having a good day, because half the assertions check the model's actual wording rather than the library's own logic. None of that is unusual, and none of it is really a "bad code" problem. It's a category mismatch; most Python packaging advice was written for libraries that call a database or parse a file, and AI libraries have a different, harder set of failure modes: outputs that aren't guaranteed to match any schema, dependencies that can be measured in gigabytes, and third-party APIs that fail in ways a normal REST client never has to think about. This article covers building robust Python AI libraries specifically — Python AI SDK best practices, and what separates a production-ready AI package from one that only survives in its own demo. Five practices, each with real, working code building toward one small, coherent example library, plus the mistakes worth watching for along the way. Prerequisites This article assumes comfort with modern Python packaging — a pyproject.toml -based project rather than a bare setup.py — working knowledge of pytest , and basic familiarity with at least one large language model or model provider's API, since the running example throughout uses one. What a Robust Python AI Library Should Actually Look Like Before trying to clear a bar, it's worth defining it clearly. A genuinely robust AI library, according to the current guidance from the Python Packaging User Guide https://packaging.python.org/tutorials/packaging-projects/ and detailed 2026 writeups on modern library construction, shares a handful of concrete traits. Full type coverage, including a py.typed marker so downstream users' type checkers actually see your annotations rather than treating your package as untyped. A pyproject.toml -first structure conforming to PEP 621 https://packaging.python.org/en/latest/specifications/pyproject-toml/ , rather than scattered configuration across several legacy files. A public API that validates its inputs and outputs rather than trusting them — which matters more here than almost anywhere else in software, since a model's output is never guaranteed to match what you asked for. Dependency isolation, so installing the library doesn't force every user into a multi-gigabyte install for a feature they'll never touch. Resilience built around every external call from day one, not bolted on after the first outage. And a continuous integration CI pipeline that enforces all of the above automatically, rather than relying on a contributor remembering to run the linter. Real Examples Worth Studying It's worth grounding all of that in real, inspectable code rather than leaving it abstract. OpenAI's own Python SDK https://github.com/openai/openai-python is worth reading for its clean top-level init .py , which exposes exactly the client, types, and exceptions a user needs and nothing more, plus a CI setup that runs both pyright and mypy against the same codebase to catch the subtle differences between type checkers. Instructor https://github.com/instructor-ai/instructor is worth studying specifically for schema-validated structured output built directly on Pydantic, which is close to the exact pattern this article's first practice walks through by hand. PydanticAI https://github.com/pydantic/pydantic-ai treats type safety as the entire design philosophy of the library rather than an afterthought bolted onto a working prototype. LiteLLM https://github.com/BerriAI/litellm is the reference point for a unified interface across dozens of providers without leaking each provider's individual quirks into the public API surface. And Hugging Face Transformers https://github.com/huggingface/transformers is the reference case, at real scale, for making heavy dependencies genuinely optional rather than bundling everything by default. The Five Best Practices, at a Glance Each of the five practices below exists to solve one specific problem that AI libraries have and ordinary libraries mostly don't: | Practice | The problem it solves | |---|---| | Schema-first public API | Model output isn't guaranteed to match what you asked for | | Testing at the large language model boundary | Model responses are nondeterministic, so naive tests flake | | Optional dependencies via extras | AI frameworks are often measured in gigabytes, not megabytes | | Resilience around external calls | Provider APIs fail in ways ordinary REST APIs rarely do | | Automated quality gates | None of the above stays true without enforcement | Practice 1: Designing a Schema-First Public API The core rule worth internalizing: never let a raw string or an untyped dictionary pulled straight from a model response cross your library's public boundary. A model call can return malformed JSON, a missing field, or a value in the wrong type, and if that raw output reaches your caller unchecked, the failure shows up somewhere far from its actual cause — usually as a confusing crash deep inside whatever code tried to use the result. Here's a real, working example: a small function that extracts structured invoice fields from raw text, validating the model's response against a Pydantic schema before it's ever allowed to leave the function. python from pydantic import BaseModel, ValidationError from openai import OpenAI client = OpenAI class ExtractedInvoice BaseModel : vendor: str total: float due date: str class SchemaValidationError Exception : """Raised when a model's response doesn't match the expected schema.""" def extract invoice raw text: str - ExtractedInvoice: """Extract structured invoice fields from raw text. Returns a validated ExtractedInvoice, never a raw dict or string.""" response = client.chat.completions.create model="gpt-4o", messages= {"role": "system", "content": "Extract invoice fields as JSON: vendor, total, due date."}, {"role": "user", "content": raw text}, , response format={"type": "json object"}, raw json = response.choices 0 .message.content try: return ExtractedInvoice.model validate json raw json except ValidationError as e: raise SchemaValidationError f"Model returned data that doesn't match ExtractedInvoice: {e}" from e Two details here matter more than they look. response format={"type": "json object"} constrains the provider to actually return valid JSON rather than prose with JSON somewhere inside it, which removes an entire class of parsing failures before validation even starts. And the except ValidationError as e: raise SchemaValidationError ... from e pattern is deliberate: it catches Pydantic's own exception type, which callers of your library shouldn't need to know or care about, and re-raises it as a clear, library-specific exception, while from e preserves the original error in the traceback for anyone who needs to debug further. The function's signature — returning ExtractedInvoice and nothing else — is itself a promise: anyone calling this function never has to write a single line of code checking whether the result "looks right." It already does. Practice 2: Testing at the Large Language Model Boundary, Not Around It This is the practice most generic Python packaging guides never cover, because it's specific to exactly this category of library. When testing custom agents, an AI-calling function actually has three testable layers: - Prompt construction - The mechanics of the call itself - How the result gets parsed The model's actual reasoning is not one of those layers — it's a black box that returns different output on different runs, and a test suite that asserts on the literal wording of a model's response will flake regardless of whether the underlying code is correct. The fix is to mock precisely at the boundary between your code and the provider, never deeper and never further out. Here's a real, runnable test suite for the extract invoice function from Practice 1: python from unittest.mock import patch, MagicMock import pytest from mylib.invoices import extract invoice, SchemaValidationError def mock response content: str - MagicMock: """Builds a fake OpenAI response object shaped just enough like the real thing for extract invoice to parse it.""" mock = MagicMock mock.choices = MagicMock message=MagicMock content=content return mock @patch "mylib.invoices.client" def test extract invoice parses valid response mock client : mock client.chat.completions.create.return value = mock response '{"vendor": "Acme Corp", "total": 452.10, "due date": "2026-09-01"}' result = extract invoice "some raw invoice text" assert result.vendor == "Acme Corp" assert result.total == 452.10 Confirm the prompt itself was constructed correctly, not just that a call happened sent messages = mock client.chat.completions.create.call args.kwargs "messages" assert "Extract invoice fields as JSON" in sent messages 0 "content" @patch "mylib.invoices.client" def test extract invoice raises on malformed output mock client : mock client.chat.completions.create.return value = mock response '{"vendor": "Acme Corp"}' missing total and due date with pytest.raises SchemaValidationError : extract invoice "some raw invoice text" @patch "mylib.invoices.client" is the single most important line in both tests: it patches the client object exactly where extract invoice looks it up — inside the mylib.invoices module — rather than patching it where it was originally defined in the openai package, which is a common and confusing mistake with unittest.mock.patch . The first test checks two genuinely different things: that the function correctly parses a well-formed response, and, separately, by inspecting call args.kwargs "messages" , that the prompt sent to the model actually contains the right instruction — catching a real class of bug where the logic runs but the wrong prompt gets sent. The second test never touches a real model at all and still verifies something true and valuable: that malformed output gets turned into a clear SchemaValidationError rather than propagating a confusing crash. Neither test depends on what a real model would say, which is exactly why they're fast, free, and won't flake in CI. Practice 3: Making Heavy Dependencies Truly Optional AI libraries have a dependency-weight problem ordinary libraries rarely face. A library offering both a hosted-API path and a local-model path shouldn't force every user into installing torch or transformers just to use the hosted path, and vice versa. The fix is pyproject.toml 's optional-dependencies extras mechanism, paired with a lazy import inside the library that fails loudly and helpfully rather than with a bare ModuleNotFoundError . project name = "mylib" dependencies = "pydantic =2.0", "httpx =0.27", project.optional-dependencies openai = "openai =1.0" local = "torch =2.0", "transformers =4.40" all = "mylib openai,local " def require module name: str, extra name: str : """Import an optional dependency, raising a clear, actionable error naming the exact extra to install if it's missing.""" try: return import module name except ImportError as e: raise ImportError f"'{module name}' is required for this feature. " f"Install it with: pip install 'mylib {extra name} '" from e def load local model model name: str : torch = require "torch", "local" transformers = require "transformers", "local" return transformers.AutoModel.from pretrained model name The core dependency list in project stays deliberately small — just pydantic and httpx , both lightweight. The project.optional-dependencies table defines named extras, so pip install mylib openai pulls in only what the hosted-API path needs, pip install mylib local pulls in the heavier local-inference stack, and pip install mylib all gets everything. The require helper is what makes this genuinely usable rather than just technically correct: without it, a user who skips the local extra and calls load local model gets a bare ModuleNotFoundError: No module named 'torch' with no indication of what to do about it. With it, they get an error that names the exact pip install command that fixes the problem — which is the difference between a five-second fix and a confused GitHub issue. Practice 4: Building Resilience Around Every External Call AI libraries live or die on the reliability of a third party they don't control. Providers rate-limit, time out, and occasionally return a 503 that clears up in a few seconds if retried, and code that doesn't account for any of that turns a routine, transient hiccup into a hard failure for every user of the library. Per a real, worked example of this pattern in Machine Learning Plus's resilient large language model client walkthrough https://machinelearningplus.com/gen-ai/resilient-llm-client/ , the fix is retry-with-backoff scoped specifically to the errors worth retrying, plus an explicit cap so a struggling provider doesn't turn into an infinite loop. python import logging import httpx from tenacity import retry, stop after attempt, wait exponential, retry if exception type, before sleep log, logger = logging.getLogger "mylib" class ProviderUnavailableError Exception : """Raised when a provider call fails after all retries are exhausted.""" @retry stop=stop after attempt 3 , hard cap, never retry forever wait=wait exponential multiplier=1, min=1, max=10 , 1s, 2s, 4s... capped at 10s retry=retry if exception type httpx.TimeoutException, httpx.HTTPStatusError , before sleep=before sleep log logger, logging.WARNING , reraise=True, on final failure, raise the real underlying error def call provider client, kwargs : return client.chat.completions.create timeout=15.0, kwargs Every parameter in that @retry decorator is doing real, deliberate work. stop after attempt 3 is the hard ceiling that stops this from ever becoming an unbounded retry loop — a genuinely dangerous failure mode where a struggling provider quietly turns into a runaway bill or a hung process. wait exponential spaces retries out with increasing delay rather than hammering an already-struggling provider immediately three times in a row. retry if exception type is what keeps this safe: it only retries on genuinely transient failures — timeouts and HTTP errors — and deliberately does not retry on, say, an authentication error, since retrying a bad API key three times wastes time without ever fixing the actual problem. before sleep log gives visibility into every retry as it happens rather than silently succeeding or failing with no trace. And reraise=True ensures that when all three attempts genuinely fail, the caller sees the real underlying exception, not a generic "retry library gave up" error that hides what actually went wrong. timeout=15.0 set explicitly on the call itself is worth noting too, since a library that never sets its own timeout is entirely at the mercy of whatever default — or lack of one — the underlying HTTP client happens to ship with. Practice 5: Automating Every Quality Gate The first four practices only stay true over time if something enforces them automatically, rather than relying on every contributor remembering to run the linter before pushing. Per the concrete, current tool stack laid out in Stephen Funk's 2026 writeup on building a Python library https://stephenlf.dev/blog/python-library-in-2026/ , the practical 2026 default is uv for environment and dependency management, ruff for both linting and formatting, mypy for type checking, and pytest with coverage for the test suite built in Practice 2 — wired together in a CI workflow that runs before anything gets released. pyproject.toml, dev dependencies dependency-groups dev = "ruff =0.6", "mypy =1.11", "pytest =8.0", "pytest-cov =5.0", "tenacity =9.0", .github/workflows/ci.yml name: CI on: push: branches: main pull request: jobs: quality: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - uses: astral-sh/setup-uv@v3 - run: uv sync --all-extras --dev - run: uv run ruff check . - run: uv run ruff format --check . - run: uv run mypy src/ - run: uv run pytest --cov=mylib --cov-report=term-missing tests/ Each step in that workflow exists to catch one specific way a contribution could quietly erode the library's reliability. ruff check and ruff format --check catch style drift and a real class of bugs ruff's linter rules flag directly, and running format --check rather than format fails the build on unformatted code instead of silently reformatting it in CI. mypy src/ verifies that the type annotations Practice 1's Pydantic schemas depend on are actually correct throughout the codebase, not just in the one function someone happened to test by hand. pytest --cov=mylib --cov-report=term-missing runs the exact boundary-mocked test suite from Practice 2 and reports which lines still aren't covered, so a gap in testing shows up as a number in the CI log instead of a surprise in production. uv sync --all-extras --dev pulls in every optional dependency from Practice 3 specifically so CI is testing the full surface of the library, not just whatever subset happens to be installed on one contributor's machine. None of these steps are exotic. What makes them a practice rather than a suggestion is that they run on every single push, automatically, whether or not anyone remembers to ask for them. Common Mistakes and Errors to Watch Out For A few mistakes come up often enough to name directly, since most of them are the exact failure mode one of the five practices above exists to prevent. - Hardcoding an API key or a specific model name directly in library code instead of accepting it as configuration, which breaks the moment anyone needs a different key or a newer model - Testing against a live provider in CI, which is slow, costly, and flaky by design — not a personal shortcoming, a structural one — exactly what Practice 2's boundary mocking exists to avoid - Trusting a model's output without validating it against a schema, the precise gap Practice 1 closes - Making a heavy framework like torch a hard, non-optional dependency for a feature most users of the library will never touch, the problem Practice 3 solves - Retrying failed calls silently and without any cap, quietly turning a temporary provider outage into a runaway bill or a hung process instead of a clear, bounded, loggable failure — exactly what Practice 4's stop after attempt guards against - Skipping the py.typed marker file, a one-line omission that quietly breaks type checking for every downstream user of an otherwise fully-typed library, since without it, type checkers treat the package as untyped no matter how carefully its own code is annotated Wrapping Up None of these five practices are really about following a checklist. They all answer the same question a library's users will eventually ask under real pressure — when a model returns something unexpected, when a provider has a bad night, when someone installs the package on a machine that can't spare four gigabytes for a dependency they don't need: can I trust this thing when something goes wrong? A library that's already answered that question before it ships, rather than after its first production incident, is the one people actually keep using. \ Shittu Olumide\ https://www.linkedin.com/in/olumide-shittu/ https://www.linkedin.com/in/olumide-shittu is a software engineer and technical writer passionate about leveraging cutting-edge technologies to craft compelling narratives, with a keen eye for detail and a knack for simplifying complex concepts. You can also find Shittu on Twitter https://twitter.com/Shittu Olumide .