{"slug": "building-a-multilingual-book-platform-with-fastapi-and-postgresql-full-text", "title": "Building a Multilingual Book Platform with FastAPI and PostgreSQL Full-Text Search", "summary": "LectuLibre, an AI-powered book translation platform, replaced slow SQL ILIKE queries with PostgreSQL full-text search to handle multilingual search across more than 10,000 books in 5 languages. The team stores one tsvector per book per language in a dedicated table with a GIN index, using PostgreSQL's language-specific text search configurations for stemming and stop words, and generates vectors asynchronously via asyncpg. The approach avoids the operational overhead of Elasticsearch or Meilisearch at their current scale.", "body_md": "*We translated 10,000+ books across 5 languages—here's how we handled multilingual search and indexing with FastAPI and PostgreSQL.*\n\nAt LectuLibre, we're building an AI-powered book translation service. Users upload an EPUB or PDF, and our platform translates it into multiple languages using LLMs like Claude and DeepSeek. While the translation pipeline is the flashy part, the backend infrastructure that serves translated books, especially search, turned out to be a significant engineering challenge.\n\nOur backend is Python/FastAPI with PostgreSQL, deployed on a VPS. As our library grew to over 10,000 books across 5 languages, we hit a wall with search performance. This post is about how we solved multilingual full-text search using PostgreSQL's built-in capabilities, and the lessons we learned along the way.\n\nInitially, we stored book metadata (title, author, description) in a simple `books` table. Search was implemented with SQL `ILIKE` queries:\n\n```\nSELECT * FROM books WHERE title ILIKE '%query%' OR author ILIKE '%query%';\n```\n\nThis worked fine for a few hundred books, but as we scaled, queries took hundreds of milliseconds—often over 500ms for a single search. Worse, it ignored language-specific nuances like stemming, stop words, and diacritics. A user searching for \"correr\" (to run in Spanish) wouldn't find books with \"corriendo\" or \"corrió\".\n\nWe needed a robust full-text search solution that:\n\nAfter evaluating options like Elasticsearch and Meilisearch, we realized that PostgreSQL's built-in full-text search could meet our needs at our current scale (10k+ books) without the operational overhead of an external service.\n\nPostgreSQL provides full-text search through `tsvector` and `tsquery` types, with built-in text search configurations for many languages. Each configuration includes a stemmer, stop word list, and parsing rules.\n\nOur plan was:\n\n`tsvector` for each book and each language it's available in.`ts_rank`.\nBecause a single book can exist in multiple languages, we created a separate table to hold language-specific search vectors.\n\nHere are the relevant tables (using SQLAlchemy models):\n\n``` python\nfrom sqlalchemy import Column, Integer, String, ForeignKey, Index\nfrom sqlalchemy.dialects.postgresql import TSVECTOR\nfrom sqlalchemy.orm import declarative_base, relationship\n\nBase = declarative_base()\n\nclass Book(Base):\n    __tablename__ = \"books\"\n    id = Column(Integer, primary_key=True)\n    original_language = Column(String(10), nullable=False)\n    # other metadata fields...\n    search_vectors = relationship(\"BookSearchVector\", back_populates=\"book\", cascade=\"all, delete-orphan\")\n\nclass BookSearchVector(Base):\n    __tablename__ = \"book_search_vectors\"\n    id = Column(Integer, primary_key=True)\n    book_id = Column(Integer, ForeignKey(\"books.id\", ondelete=\"CASCADE\"), nullable=False)\n    language = Column(String(10), nullable=False)\n    vector = Column(TSVECTOR, nullable=False)\n    book = relationship(\"Book\", back_populates=\"search_vectors\")\n\n    __table_args__ = (\n        Index(\"ix_book_search_vectors_language_vector\", \"language\", \"vector\", postgresql_using=\"gin\"),\n    )\n```\n\nWe chose to store one `tsvector` per language rather than a combined vector. This allows us to search specifically in one language or across languages by querying multiple rows.\n\nWhen a book is added or a new translation is completed, we generate the `tsvector` using PostgreSQL's `to_tsvector` with the appropriate language configuration. Since our backend is async, we used `asyncpg` directly for efficiency:\n\n``` python\nimport asyncpg\n\nasync def update_search_vector(book_id: int, language: str, title: str, author: str, description: str):\n    conn = await asyncpg.connect(DATABASE_URL)\n    try:\n        await conn.execute(\n            \"\"\"\n            INSERT INTO book_search_vectors (book_id, language, vector)\n            VALUES ($1, $2, to_tsvector($3::regconfig, $4 || ' ' || $5 || ' ' || $6))\n            ON CONFLICT (book_id, language) DO UPDATE\n            SET vector = EXCLUDED.vector\n            \"\"\",\n            book_id,\n            language,\n            language,  # e.g., 'english', 'spanish', 'french'\n            title,\n            author,\n            description\n        )\n    finally:\n        await conn.close()\n```\n\nNote: We didn't have a unique constraint on `(book_id, language)` initially, leading to duplicate rows. We added one after discovering the issue.\n\nFor languages not supported by PostgreSQL's built-in configurations (like some regional variants), we fall back to the `'simple'` configuration, which lowercases and splits words but doesn't stem or remove stop words. We also applied the `unaccent` extension to remove diacritics, making searches more forgiving.\n\nOur search endpoint accepts a query string and an optional language filter. If no language is specified, we search across all languages and merge results.\n\n``` python\nfrom fastapi import FastAPI, Query\nfrom sqlalchemy import text\nfrom sqlalchemy.ext.asyncio import AsyncSession\n\napp = FastAPI()\n\n@app.get(\"/search\")\nasync def search(\n    q: str = Query(..., min_length=2),\n    lang: str | None = Query(None, regex=\"^[a-z]{2,3}$\"),\n    db: AsyncSession = Depends(get_db)\n):\n    # Build the query\n    if lang:\n        tsquery = f\"websearch_to_tsquery('{lang}', :q)\"\n        language_filter = f\"AND language = '{lang}'\"\n    else:\n        # Use 'simple' for cross-language search (no stemming)\n        tsquery = \"websearch_to_tsquery('simple', :q)\"\n        language_filter = \"\"\n\n    sql = f\"\"\"\n        SELECT b.id, b.title, b.author, b.original_language,\n               ts_rank(sv.vector, {tsquery}) AS rank\n        FROM book_search_vectors sv\n        JOIN books b ON b.id = sv.book_id\n        WHERE sv.vector @@ {tsquery}\n        {language_filter}\n        ORDER BY rank DESC\n        LIMIT 20\n    \"\"\"\n    result = await db.execute(text(sql), {\"q\": q})\n    return result.fetchall()\n```\n\nWe used `websearch_to_tsquery` because it handles user-friendly syntax (like Google search) and automatically adds `&` between words. For language-specific searches, we pass the language code directly.\n\nBefore optimization, our `ILIKE` search took **500-800ms** on average for 10k books. After implementing `tsvector` with GIN index, the same searches now complete in **5-15ms**—a 50-100x improvement. The index adds about 20% overhead to write operations, which is acceptable since search reads far outnumber writes.\n\nWe also monitored PostgreSQL memory settings. Initially, GIN index scans were slow due to low `work_mem`. Increasing it from 4MB to 16MB improved indexing speed by 30%.\n\nPostgreSQL ships with configurations for about 30 languages, but not all dialects are covered. For example, we had to create a custom configuration for Latin American Spanish by copying the `spanish` config and adjusting stop words. For languages without built-in support, we used `simple` and accepted less optimal stemming.\n\nWe initially tried database triggers to update `tsvector` automatically on insert/update. However, our async workflow made it tricky to manage, and we often ended up with stale data during translation updates. We moved to updating vectors in Python after translation completion, which gave us more control and easier debugging.\n\nWhen a user searches without specifying a language, we need to search across all languages. Using the `'simple'` configuration avoids stemming but still provides decent results. However, for better relevance, we could combine results from multiple language-specific searches, but we haven't needed that yet.\n\nAt our scale (10k books, <1M search vectors), PostgreSQL is more than sufficient. We saved ourselves the operational burden of running an Elasticsearch cluster. We may revisit this decision if we grow 10x.\n\n`work_mem` and `maintenance_work_mem`.\nBuilding a multilingual platform with FastAPI and PostgreSQL has been a rewarding journey. Leveraging PostgreSQL's full-text search allowed us to deliver fast, language-aware search without adding complexity. The code examples above are simplified but should give you a starting point.\n\n**Open question for the community:** How do you handle full-text search for languages not supported by your database's built-in configurations? We'd love to hear your approaches.\n\nIf you're building a similar system, start with PostgreSQL full-text search—it might be all you need.", "url": "https://wpnews.pro/news/building-a-multilingual-book-platform-with-fastapi-and-postgresql-full-text", "canonical_source": "https://dev.to/jacob_gong/building-a-multilingual-book-platform-with-fastapi-and-postgresql-full-text-search-3gce", "published_at": "2026-09-30 03:01:47+00:00", "updated_at": "2026-09-30 03:17:06.927133+00:00", "lang": "en", "topics": ["ai-products", "developer-tools"], "entities": ["LectuLibre", "FastAPI", "PostgreSQL", "Elasticsearch", "Meilisearch", "asyncpg", "Claude", "DeepSeek"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/building-a-multilingual-book-platform-with-fastapi-and-postgresql-full-text", "markdown": "https://wpnews.pro/news/building-a-multilingual-book-platform-with-fastapi-and-postgresql-full-text.md", "text": "https://wpnews.pro/news/building-a-multilingual-book-platform-with-fastapi-and-postgresql-full-text.txt", "jsonld": "https://wpnews.pro/news/building-a-multilingual-book-platform-with-fastapi-and-postgresql-full-text.jsonld"}}