Build a documentation chatbot with LangChain, OpenAI, Pinecone, and Apify A tutorial published by Apify details how to build a documentation chatbot using LangChain, OpenAI, Pinecone, and Apify, with UpCloud's documentation as the example corpus and Streamlit for the chat interface. The guide walks through collecting documentation via Apify, splitting it into passages stored in Pinecone with source URLs, generating OpenAI embeddings, and wiring retrieval-augmented generation (RAG) through LangChain with LangGraph retaining conversation history for follow-ups. It requires Python 3.12 plus OpenAI, Pinecone, and Apify accounts, and notes that a ChatGPT subscription does not cover OpenAI API usage. You can spend more time searching documentation than completing the task that brought you there. One guide assumes you already know the terminology. Another covers only part of the setup. When the instructions don’t quite fit your situation, you’re back to searching and piecing things together yourself. A documentation chatbot can make that process easier. Users can describe what they need in plain language, get guidance that fits their setup, and ask follow-up questions as their situation changes. In this tutorial, you’ll build one with LangChain, OpenAI, Pinecone, and Apify, using UpCloud’s documentation https://upcloud.com/docs/ as the example. You’ll collect its user guides, make them searchable for AI, and build a chat interface with Streamlit that you can publish and share for free. Test the finished UpCloud Docs Companion , then clone or download the complete project from GitHub . Follow the steps below to run it with your own API keys. How the documentation chatbot works UpCloud is a cloud platform where you can host applications and databases. Suppose you want to connect a Python app to its Managed PostgreSQL database service without making the database publicly accessible. The chatbot searches pre-indexed documentation and explains the setup, with source links you can verify. If you then ask, “What changes if I’m developing on my laptop?”, it uses the conversation history to understand which setup you mean and which requirements still apply. To make that possible, the chatbot app first collects documentation through Apify, splits it into passages, and stores those passages in Pinecone alongside their source URLs. OpenAI turns the text into embeddings https://aws.amazon.com/what-is/embeddings-in-machine-learning/ , numerical representations that help search match related meanings. A question about keeping a database off the internet can therefore find guidance about private networking, even when the wording differs. Using retrieved material to help generate an answer is called retrieval-augmented generation RAG https://aws.amazon.com/what-is/retrieval-augmented-generation/ . LangChain connects the chatbot’s large language model https://www.signitysolutions.com/blog/top-large-language-models to the documentation search and a second tool for live lookups through Apify. When the saved collection leaves a gap, the assistant can retrieve additional pages. LangGraph keeps earlier messages available for follow-ups, while Streamlit hosts and presents the conversation as an interactive UI. Prerequisites You’ll need Python 3.12 https://www.python.org/downloads/ , a code editor such as PyCharm https://www.jetbrains.com/pycharm/download/ , and accounts with OpenAI https://platform.openai.com/signup , Pinecone https://app.pinecone.io/?sessionType=login , and Apify https://console.apify.com/sign-up . Make sure your OpenAI account has API billing enabled and your Apify account has available credits. A ChatGPT subscription doesn’t cover OpenAI API usage. To publish the app, you’ll also need GitHub https://github.com/signup and Streamlit Community Cloud https://streamlit.io/cloud accounts. Community Cloud provides free hosting, but OpenAI, Pinecone, and Apify have their own usage limits and charges. Step 1: Set up the Python project Start by creating a project in PyCharm or any editor of your choice with its own virtual environment. This keeps the chatbot’s packages separate from those used in your other projects. Create the project in PyCharm - Select New Project → Pure Python . - Name the project folder upcloud-docs-bot . - Select Python 3.12 as the interpreter. - Create a project virtual environment named .venv . - Click Create . Add the project files Create the following files directly inside upcloud-docs-bot : - ingest.py : Collects the documentation and indexes it for search. - app.py : Runs the chatbot and its interface. - requirements.txt : Lists the Python packages the project needs. - .env : Stores your local API credentials. - .gitignore : Tells Git which files to exclude, including credentials and the virtual environment. Use New → Python File for ingest.py and app.py . For the remaining files, use New → File . namespace.txt will only appear in your directory after you run ingest.py, so ignore mine displayed above for now. List the dependencies Paste the following into requirements.txt and save it: langchain==1.4.2 langchain-openai==1.6.2 langchain-pinecone==0.2.13 langchain-apify==0.1.7 langchain-text-splitters==1.1.2 apify-client==2.5.1 streamlit==1.64.0 python-dotenv==1.2.3 The LangChain packages connect OpenAI, Pinecone, and Apify and split documentation into passages for search. They also install LangGraph, which manages conversation history. The apify-client package runs the crawler, python-dotenv loads credentials from .env , and streamlit provides the chat interface. Install the dependencies Open PyCharm’s Terminal in the upcloud-docs-bot folder and run this command for your operating system. On Windows : .\.venv\Scripts\python.exe -m pip install -r requirements.txt On macOS or Linux : .venv/bin/python -m pip install -r requirements.txt Both commands use Python from the project’s virtual environment, so you don’t need to activate it first. Step 2: Add API credentials and create a Pinecone index Add your API credentials so the app can access each service. Then create a Pinecone index to store and search the documentation’s embeddings: Get your API credentials - For OpenAI: Create a key on the API keys page https://platform.openai.com/api-keys . - Pinecone: Open the Pinecone console https://app.pinecone.io/ ; as soon as you sign in, you’ll be given a default API key . - Apify: Find your API token in Apify’s integrations settings https://console.apify.com/settings/integrations . Open .env and paste the following: OPENAI API KEY="your openai key" PINECONE API KEY="your pinecone key" APIFY TOKEN="your apify token" PINECONE INDEX NAME="upcloud-docs" Replace the three credential placeholders with your keys and token. Keep upcloud-docs as the index name; you’ll create that index below. Keep your credentials and local files out of Git Add these rules to .gitignore to exclude your credentials, virtual environment, editor settings, and temporary files: .env .venv/ .idea/ pycache / namespace.tmp .streamlit/secrets.toml Create the Pinecone index In the Pinecone console, open Database → Indexes → Create index . Under Configuration , scroll to the right and select OpenAI’s text-embedding-3-small model. Set Dimension to 1536 and make sure your settings match the following: | Setting | Value | |---|---| | Name | upcloud-docs | | Configuration | OpenAI text-embedding-3-small | | Vector type | Dense | | Dimension | 1536 | | Metric | cosine | | Deployment | Serverless | | Cloud and region | AWS, us-east-1 | Both scripts call OpenAI to generate the embeddings, while Pinecone stores and searches those vectors. The 1536 dimension setting matches the model’s default embedding size, keeping the index compatible with the vectors your scripts produce. Click Create index and wait until the index is ready before continuing. Step 3: Collect and index the documentation Now build the documentation index your chatbot will search. The script uses Website Content Crawler https://apify.com/apify/website-content-crawler , an Apify Actor, to collect UpCloud’s Managed PostgreSQL and networking guides, along with two PostgreSQL reference pages covering connections and SSL. Add the ingestion script Open ingest.py paste in the complete script from the project repository https://github.com/Gogo-Egop/upcloud-docs-bot , and save the file. The excerpt below only shows how the script starts the crawl: python import os from datetime import datetime, timezone from pathlib import Path from uuid import uuid4 from apify client import ApifyClient from dotenv import load dotenv from langchain core.documents import Document from langchain openai import OpenAIEmbeddings from langchain pinecone import PineconeVectorStore from langchain text splitters import RecursiveCharacterTextSplitter ROOT = Path file .resolve .parent load dotenv ROOT / ".env" UPCLOUD = "https://upcloud.com/docs/products/managed-postgresql/", "https://upcloud.com/docs/products/networking/", POSTGRES = "https://www.postgresql.org/docs/18/libpq-connect.html", "https://www.postgresql.org/docs/18/libpq-ssl.html", client = ApifyClient os.environ "APIFY TOKEN" print "Collecting documentation with Apify..." run = client.actor "apify/website-content-crawler" .call run input={ "startUrls": {"url": url} for url in UPCLOUD + POSTGRES , "includeUrlGlobs": url + " " for url in UPCLOUD , "crawlerType": "playwright:adaptive", "maxCrawlPages": 100, "maxCrawlDepth": 5, "maxConcurrency": 3, "saveMarkdown": True, "respectRobotsTxtFile": True, "proxyConfiguration": {"useApifyProxy": True}, }, timeout secs=1800, if not run or run "status" = "SUCCEEDED": raise RuntimeError "Crawl did not finish. Check the run in Apify Console." Excerpt from ingest.py; use the complete script from: https://github.com/Gogo-Egop/upcloud-docs-bot How the script prepares the documentation The script starts by collecting the pages. startUrls tells the crawler where to begin, while includeUrlGlobs restricts the links it follows to the two UpCloud documentation sections. It supplies the two PostgreSQL reference pages directly as starting URLs and caps the collection at 100 pages and a crawl depth of five. Once the crawl completes, the script filters out pages outside the selected sources, pages with HTTP errors, and pages with fewer than 100 characters of content. It turns each accepted page into a LangChain Document , keeping its title, text, source URL, and collection time together. Next, the script splits long pages into passages of up to 3,000 characters. The splitter uses a 300-character overlap to preserve context between neighboring passages. OpenAI then generates embeddings for these passages, and the script uploads them to a new Pinecone namespace , which is a named group of records within the index. After the uploads, the script saves the namespace’s name in namespace.txt so the chatbot knows which collection to search. Run the ingestion script Open PyCharm’s Terminal in your project folder. On Windows, run: .\.venv\Scripts\python.exe ingest.py On macOS or Linux, run: .venv/bin/python ingest.py Confirm that the indexing ran completely Wait for the terminal to show the page and passage counts, followed by a message that says Ready. Run: python -m streamlit run app.py . Then check that namespace.txt appears in your project folder. You can also open the crawl under Runs in Apify Console https://console.apify.com/ to inspect the collected pages. Step 4: Build the chatbot Now connect your documentation index to the AI assistant and give readers a chat interface to query from. To follow along, open app.py in the project repository https://github.com/Gogo-Egop/upcloud-docs-bot to review or copy the complete script. The sections below walk through the code in order, using short excerpts to explain the main parts. Load the libraries and credentials The imports provide the chat model, search tools, Streamlit interface, and conversation memory. After those imports, the script sets the project path and loads your local API settings: ROOT = Path file .resolve .parent load dotenv ROOT / ".env" ROOT points to the folder containing app.py , so the app can find .env and namespace.txt . The load dotenv call reads your settings from .env . Define how the assistant should answer The SYSTEM PROMPT gives the assistant instructions for researching questions and writing answers. It tells the AI assistant to search the indexed documentation, look up missing details, and cite sources returned by its tools. It also asks the assistant to keep the reader’s goal and constraints in context as they add or change details. For database questions, the prompt requires verified connection settings and clear explanations of any prerequisites. It distinguishes private routing , which determines how the application reaches the database, from certificate verification , which checks the server’s identity. The prompt also contains a SQL example that the assistant can include in its answers. That query reports whether a connection uses TLS, along with its protocol and cipher. The chatbot doesn't execute it. Keep the full prompt from the repository in app.py , including these instructions. Handle source checks and long pages The next three helpers handle explicit requests to check a source and details buried in long reference pages. current turn retrieves the latest user question and every message that follows it, including tool responses. This gives the app a complete view of what has happened so far in the current turn. require requested lookup detects common user requests, like “check the official docs.” If the live lookup tool is available, it requires the assistant to try it before answering. When the wording falls outside the helper’s recognized patterns, the model decides whether a lookup is needed. select excerpt controls how much page content reaches the model. It returns pages of up to 12,000 characters in full. For longer pages, it selects four passages with the most matches for the supplied focus words. For example, “account CA certificate” helps it find certificate instructions within a long API reference. CA means "certificate authority". Connect documentation search and live lookups The build agent function connects the assistant to two tools. search docs searches Pinecone for up to five related passages and returns their source URLs. It uses the same embedding model as ingest.py , so questions and indexed passages can be compared. lookup live docs uses Apify’s LangChain integration https://docs.apify.com/integrations/langchain to search the web or open a URL. It checks whether the run succeeded and filters unusable results before returning excerpts. Both functions use the @tool decorator, which makes them available to the model. At the end of build agent , create agent https://docs.langchain.com/oss/python/langchain/agents brings the model, tools, system prompt, and memory together: return create agent model=ChatOpenAI model="gpt-5.4-mini", timeout=90, max retries=2 , tools= search docs, lookup live docs , checkpointer=InMemorySaver , middleware= require requested lookup, ToolCallLimitMiddleware run limit=6 , ToolCallLimitMiddleware tool name="lookup live docs", run limit=3 , , system prompt=system prompt, The model can call a tool, read its results, and decide whether it needs more information before answering. The tool limits allow up to six calls per turn, including at most three live lookups. The @st.cache resource decorator on build agent lets Streamlit reuse the agent when the page reruns. Live results remain in the conversation without being added to Pinecone. The system prompt asks the model to prefer official sources, but the browsing tool does not enforce a domain restriction. Let readers inspect the sources The Documentation checked panel beneath each answer lets readers see which tools ran, whether they returned usable content, and which URLs they retrieved. collect checks reads the current turn’s tool responses and extracts their status, source URLs, and any error type. show checks displays those details in a collapsible panel. Readers can use the panel alongside the answer’s citations to inspect the evidence. They still need to check whether the returned pages support the instructions. Add the chat interface and conversation memory The final section of app.py builds the Streamlit interface. Before accepting questions, it checks for the required API settings and namespace.txt . If anything is missing, it explains what to add or run. Each conversation gets a random thread id . When the reader submits a question, the app passes that question and the thread ID to the agent: result = agent.invoke {"messages": {"role": "user", "content": question} }, config={ "configurable": {"thread id": st.session state.thread id}, "recursion limit": 30, }, LangGraph uses the thread ID to retrieve earlier messages, so agent.invoke only needs the new question. Separately, Streamlit’s session state keeps previous messages and source panels visible when the page reruns. Clicking New conversation creates a new thread ID and clears the visible chat. Conversation memory is temporary: a server restart clears it, and a new browser session starts a new chat. Step 5: Run the app and try a conversation You can now test whether the chatbot functions as intended, at least at the fundamental level: searches the documentation, retrieves additional sources, and remembers your setup. Start the app Save app.py Ctrl + S and open PyCharm’s Terminal in the project folder. On Windows, run: .\.venv\Scripts\python.exe -m streamlit run app.py On macOS or Linux, run: .venv/bin/python -m streamlit run app.py Open the local address shown in the terminal, usually http://localhost:8501 . Keep the terminal running while you use the chatbot. Test the documentation search The first question gives the chatbot both a task and a constraint: I have a Python app on an UpCloud server. I want to connect it privately to UpCloud Managed PostgreSQL. Explain what I need to set up. In the response below, the chatbot explains the private network route and connection setup. The expanded Documentation checked panel shows the indexed pages it retrieved, with links to the sources used in the answer. Test an external lookup The next question asks the chatbot to consult documentation outside the indexed collection: I’m using Psycopg 3. Check the official driver documentation and show me the connection settings, including certificate verification. Here, the Documentation checked panel shows a live lookup through Apify. The chatbot uses the retrieved Psycopg documentation to explain the connection settings, including how to configure certificate verification. Test conversation memory The final question changes where the application runs without repeating the earlier setup: Actually, during development the app runs on my laptop. What changes? The response below carries forward the earlier context: UpCloud Managed PostgreSQL, Psycopg 3, and the requirement to keep the database private. It focuses on what changes when the application runs on a laptop, explaining how to reach the private network while preserving certificate verification. Step 6: Publish the app with Streamlit Once the chatbot works locally, you can publish it so others can use it in their browsers. Here's how to set it up: Upload the project to GitHub - Create a GitHub repository https://github.com/new named upcloud-docs-bot and initialize it with a README. - Select Add file → Upload files . - Upload app.py , ingest.py , requirements.txt , namespace.txt , and .gitignore to the repository’s main folder, alongside the README. - Commit the files. Use the namespace.txt generated by ingest.py . It tells the hosted app which collection to search in Pinecone. Keep .env and .venv off GitHub. When uploading through the website, .gitignore does not filter the files you select. Configure the hosted app Open Streamlit Community Cloud https://share.streamlit.io/ , connect your GitHub account, and choose Create app . Enter the repository details: | Field | Value | |---|---| | Repository | YOUR GITHUB USERNAME/upcloud-docs-bot | | Branch | Your repository’s branch, usually main | | Main file path | app.py | Add the API credentials Open Advanced settings and select Python 3.12 . Paste the following into Secrets , replacing the three credential placeholders: OPENAI API KEY = "your openai key" PINECONE API KEY = "your pinecone key" APIFY TOKEN = "your apify token" PINECONE INDEX NAME = "upcloud-docs" Keep these settings at the top level, without a section heading. Streamlit makes root-level secrets available as environment variables https://docs.streamlit.io/develop/concepts/connections/secrets-management , so the app can read them the same way it reads your local settings. Deploy and share - Save the settings and deploy the app https://docs.streamlit.io/deploy/streamlit-community-cloud/deploy-your-app/deploy . Streamlit will install the dependencies and run app.py . - Open the hosted app and repeat the conversation checks from Step 5. - Under Settings → Sharing , choose who can access the app and copy its URL. Visitors’ questions use your API accounts, so monitor usage when sharing the app, especially publicly. The indexed collection reflects your last import. To refresh it: - Run ingest.py again to collect and index the documentation. - Upload the newly generated namespace.txt to GitHub, replacing the previous version. - Confirm that the hosted chatbot can search the new collection. - Remove older Pinecone namespaces you no longer need after confirming the new collection works. Adapt this chatbot to your documentation Finding the right guide is only half the work. The harder part is connecting scattered instructions to your setup and keeping that context as the problem evolves. UpCloud Docs Companion brings those sources into one conversation, lets readers check the guidance, and handles follow-up questions without making them start over or repeat themselves. To adapt this approach to your documentation, update the starting URLs, crawl filters, coverage checks, and system prompt. Make sure the index dimensions still match the embedding model. Create a free Apify account https://console.apify.com/sign-up and use the included $5 in free credit https://apify.com/pricing to test your first crawl. Explore thousands of ready-to-run Actors on Apify Store https://apify.com/store , including scrapers built for specific websites, to add new sources to your chatbot. Contact sales https://apify.com/contact-sales