# Build a documentation chatbot with LangChain, OpenAI, Pinecone, and Apify

> Source: <https://blog.apify.com/how-to-use-langchain/>
> Published: 2026-09-29 06:25:00+00:00

You can spend more time searching documentation than completing the task that brought you there.

One guide assumes you already know the terminology. Another covers only part of the setup. When the instructions don’t quite fit your situation, you’re back to searching and piecing things together yourself.

A documentation chatbot can make that process easier. Users can describe what they need in plain language, get guidance that fits their setup, and ask follow-up questions as their situation changes.

In this tutorial, you’ll build one with LangChain, OpenAI, Pinecone, and Apify, using [UpCloud’s documentation](https://upcloud.com/docs/) as the example. You’ll collect its user guides, make them searchable for AI, and build a chat interface with Streamlit that you can publish and share for free.

*Test the finished*

*UpCloud Docs Companion*

*, then clone or download the*

*complete project from GitHub*

*. Follow the steps below to run it with your own API keys.*
## How the documentation chatbot works

UpCloud is a cloud platform where you can host applications and databases. Suppose you want to connect a Python app to its Managed PostgreSQL database service without making the database publicly accessible. The chatbot searches pre-indexed documentation and explains the setup, with source links you can verify.

If you then ask, “What changes if I’m developing on my laptop?”, it uses the conversation history to understand which setup you mean and which requirements still apply.

To make that possible, the chatbot app first collects documentation through Apify, splits it into passages, and stores those passages in Pinecone alongside their source URLs. OpenAI turns the text into [**embeddings**](https://aws.amazon.com/what-is/embeddings-in-machine-learning/), numerical representations that help search match related meanings.

A question about keeping a database off the internet can therefore find guidance about private networking, even when the wording differs. Using retrieved material to help generate an answer is called [retrieval-augmented generation (RAG](https://aws.amazon.com/what-is/retrieval-augmented-generation/)).

LangChain connects the chatbot’s [large language model](https://www.signitysolutions.com/blog/top-large-language-models) to the documentation search and a second tool for live lookups through Apify. When the saved collection leaves a gap, the assistant can retrieve additional pages.

LangGraph keeps earlier messages available for follow-ups, while Streamlit hosts and presents the conversation as an interactive UI.

## Prerequisites

You’ll need [Python 3.12](https://www.python.org/downloads/), a code editor such as [PyCharm](https://www.jetbrains.com/pycharm/download/), and accounts with [OpenAI](https://platform.openai.com/signup), [Pinecone](https://app.pinecone.io/?sessionType=login), and [Apify](https://console.apify.com/sign-up). Make sure your OpenAI account has API billing enabled and your Apify account has available credits. A ChatGPT subscription doesn’t cover OpenAI API usage.

To publish the app, you’ll also need [GitHub](https://github.com/signup) and [Streamlit Community Cloud](https://streamlit.io/cloud) accounts. Community Cloud provides free hosting, but OpenAI, Pinecone, and Apify have their own usage limits and charges.

## Step 1: Set up the Python project

Start by creating a project in PyCharm (or any editor of your choice) with its own virtual environment. This keeps the chatbot’s packages separate from those used in your other projects.

### Create the project in PyCharm

- Select **New Project → Pure Python** .
- Name the project folder `upcloud-docs-bot` .
- Select **Python 3.12** as the interpreter.
- Create a project virtual environment named `.venv` .
- Click **Create** .

### Add the project files

Create the following files directly inside `upcloud-docs-bot`:

- `ingest.py` : Collects the documentation and indexes it for search.
- `app.py` : Runs the chatbot and its interface.
- `requirements.txt` : Lists the Python packages the project needs.
- `.env` : Stores your local API credentials.
- `.gitignore` : Tells Git which files to exclude, including credentials and the virtual environment.

Use **New → Python File** for `ingest.py` and `app.py`. For the remaining files, use **New → File**.

*namespace.txt will only appear in your directory after you run ingest.py, so ignore mine displayed above for now.*
### List the dependencies

Paste the following into `requirements.txt` and save it:

```
langchain==1.4.2
langchain-openai==1.6.2
langchain-pinecone==0.2.13
langchain-apify==0.1.7
langchain-text-splitters==1.1.2
apify-client==2.5.1
streamlit==1.64.0
python-dotenv==1.2.3
```

The LangChain packages connect OpenAI, Pinecone, and Apify and split documentation into passages for search. They also install LangGraph, which manages conversation history.

The`apify-client` package runs the crawler, `python-dotenv` loads credentials from `.env`, and` streamlit` provides the chat interface.

### Install the dependencies

Open PyCharm’s **Terminal** in the `upcloud-docs-bot` folder and run this command for your operating system.

On **Windows**:

```
.\.venv\Scripts\python.exe -m pip install -r requirements.txt
```

On **macOS or Linux**:

```
.venv/bin/python -m pip install -r requirements.txt
```

Both commands use Python from the project’s virtual environment, so you don’t need to activate it first.

## Step 2: Add API credentials and create a Pinecone index

Add your API credentials so the app can access each service. Then create a Pinecone index to store and search the documentation’s embeddings:

### Get your API credentials

- **For OpenAI:** Create a key on the[API keys page](https://platform.openai.com/api-keys) .
- **Pinecone:** Open the[Pinecone console](https://app.pinecone.io/) ; as soon as you sign in, you’ll be given a default**API key** .
- **Apify:** Find your API token in[Apify’s integrations settings](https://console.apify.com/settings/integrations) .

Open `.env` and paste the following:

```
OPENAI_API_KEY="your_openai_key"
PINECONE_API_KEY="your_pinecone_key"
APIFY_TOKEN="your_apify_token"
PINECONE_INDEX_NAME="upcloud-docs"
```

Replace the three credential placeholders with your keys and token. Keep `upcloud-docs` as the index name; you’ll create that index below.

### Keep your credentials and local files out of Git

Add these rules to `.gitignore` to exclude your credentials, virtual environment, editor settings, and temporary files:

```
.env
.venv/
.idea/
__pycache__/
namespace.tmp
.streamlit/secrets.toml
```

### Create the Pinecone index

In the Pinecone console, open **Database → Indexes → Create index**. Under **Configuration**, scroll to the right and select OpenAI’s **text-embedding-3-small** model. Set **Dimension** to **1536** and make sure your settings match the following:

| Setting | Value | 
|---|---|
| Name | `upcloud-docs` | 
| Configuration | OpenAI `text-embedding-3-small` | 
| Vector type | Dense | 
| Dimension | `1536` | 
| Metric | `cosine` | 
| Deployment | Serverless | 
| Cloud and region | AWS, `us-east-1` | 

Both scripts call OpenAI to generate the embeddings, while Pinecone stores and searches those vectors. The **1536** dimension setting matches the model’s default embedding size, keeping the index compatible with the vectors your scripts produce.

Click **Create index** and wait until the index is ready before continuing.

## Step 3: Collect and index the documentation

Now build the documentation index your chatbot will search. The script uses [Website Content Crawler](https://apify.com/apify/website-content-crawler), an Apify Actor, to collect UpCloud’s Managed PostgreSQL and networking guides, along with two PostgreSQL reference pages covering connections and SSL.

### Add the ingestion script

Open `ingest.py` paste in the complete script from the [project repository](https://github.com/Gogo-Egop/upcloud-docs-bot), and save the file. The excerpt below only shows how the script starts the crawl:

``` python
import os
from datetime import datetime, timezone
from pathlib import Path
from uuid import uuid4

from apify_client import ApifyClient
from dotenv import load_dotenv
from langchain_core.documents import Document
from langchain_openai import OpenAIEmbeddings
from langchain_pinecone import PineconeVectorStore
from langchain_text_splitters import RecursiveCharacterTextSplitter

ROOT = Path(__file__).resolve().parent
load_dotenv(ROOT / ".env")

UPCLOUD = [
    "https://upcloud.com/docs/products/managed-postgresql/",
    "https://upcloud.com/docs/products/networking/",
]
POSTGRES = [
    "https://www.postgresql.org/docs/18/libpq-connect.html",
    "https://www.postgresql.org/docs/18/libpq-ssl.html",
]

client = ApifyClient(os.environ["APIFY_TOKEN"])
print("Collecting documentation with Apify...")
run = client.actor("apify/website-content-crawler").call(
    run_input={
        "startUrls": [{"url": url} for url in UPCLOUD + POSTGRES],
        "includeUrlGlobs": [url + "**" for url in UPCLOUD],
        "crawlerType": "playwright:adaptive",
        "maxCrawlPages": 100,
        "maxCrawlDepth": 5,
        "maxConcurrency": 3,
        "saveMarkdown": True,
        "respectRobotsTxtFile": True,
        "proxyConfiguration": {"useApifyProxy": True},
    },
    timeout_secs=1800,
)
if not run or run["status"] != "SUCCEEDED":
    raise RuntimeError("Crawl did not finish. Check the run in Apify Console.")

# Excerpt from ingest.py; use the complete script from:
# https://github.com/Gogo-Egop/upcloud-docs-bot
```

### How the script prepares the documentation

The script starts by collecting the pages. `startUrls` tells the crawler where to begin, while `includeUrlGlobs` restricts the links it follows to the two UpCloud documentation sections. It supplies the two PostgreSQL reference pages directly as starting URLs and caps the collection at 100 pages and a crawl depth of five.

Once the crawl completes, the script filters out pages outside the selected sources, pages with HTTP errors, and pages with fewer than 100 characters of content. It turns each accepted page into a LangChain `Document`, keeping its title, text, source URL, and collection time together.

Next, the script splits long pages into passages of up to 3,000 characters. The splitter uses a 300-character overlap to preserve context between neighboring passages.

OpenAI then generates embeddings for these passages, and the script uploads them to a new Pinecone **namespace**, which is a named group of records within the index. After the uploads, the script saves the namespace’s name in `namespace.txt` so the chatbot knows which collection to search.

### Run the ingestion script

Open PyCharm’s **Terminal** in your project folder. On Windows, run:

```
.\.venv\Scripts\python.exe ingest.py
```

On macOS or Linux, run:

```
.venv/bin/python ingest.py
```

### Confirm that the indexing ran completely 

Wait for the terminal to show the page and passage counts, followed by a message that says `Ready. Run: python -m streamlit run app.py`. Then check that `namespace.txt` appears in your project folder. You can also open the crawl under **Runs** in [Apify Console](https://console.apify.com/) to inspect the collected pages.

## Step 4: Build the chatbot

Now connect your documentation index to the AI assistant and give readers a chat interface to query from. To follow along, open `app.py` in the [project repository](https://github.com/Gogo-Egop/upcloud-docs-bot) to review or copy the complete script. The sections below walk through the code in order, using short excerpts to explain the main parts.

### Load the libraries and credentials

The imports provide the chat model, search tools, Streamlit interface, and conversation memory. After those imports, the script sets the project path and loads your local API settings:

```
ROOT = Path(__file__).resolve().parent
load_dotenv(ROOT / ".env")
```

`ROOT` points to the folder containing `app.py`, so the app can find `.env` and `namespace.txt`. The `load_dotenv()` call reads your settings from `.env`.

### Define how the assistant should answer

The `SYSTEM_PROMPT` gives the assistant instructions for researching questions and writing answers. It tells the AI assistant to search the indexed documentation, look up missing details, and cite sources returned by its tools. It also asks the assistant to keep the reader’s goal and constraints in context as they add or change details.

For database questions, the prompt requires verified connection settings and clear explanations of any prerequisites. It distinguishes **private routing**, which determines how the application reaches the database, from **certificate verification**, which checks the server’s identity.

The prompt also contains a SQL example that the assistant can include in its answers. That query reports whether a connection uses TLS, along with its protocol and cipher. The chatbot doesn't execute it.

Keep the full prompt from the repository in `app.py`, including these instructions.

### Handle source checks and long pages

The next three helpers handle explicit requests to check a source and details buried in long reference pages.

`current_turn()` retrieves the latest user question and every message that follows it, including tool responses. This gives the app a complete view of what has happened so far in the current turn.

`require_requested_lookup()` detects common user requests, like “check the official docs.” If the live lookup tool is available, it requires the assistant to try it before answering. When the wording falls outside the helper’s recognized patterns, the model decides whether a lookup is needed. 

`select_excerpt()` controls how much page content reaches the model. It returns pages of up to 12,000 characters in full. For longer pages, it selects four passages with the most matches for the supplied focus words. For example, “account CA certificate” helps it find certificate instructions within a long API reference. 

*CA means "certificate authority".*
### Connect documentation search and live lookups

The `build_agent()` function connects the assistant to two tools.

`search_docs()` searches Pinecone for up to five related passages and returns their source URLs. It uses the same embedding model as `ingest.py`, so questions and indexed passages can be compared.

`lookup_live_docs()` uses [Apify’s LangChain integration](https://docs.apify.com/integrations/langchain) to search the web or open a URL. It checks whether the run succeeded and filters unusable results before returning excerpts.

Both functions use the `@tool` decorator, which makes them available to the model. At the end of `build_agent()`, [` create_agent()`](https://docs.langchain.com/oss/python/langchain/agents) brings the model, tools, system prompt, and memory together:

```
return create_agent(
    model=ChatOpenAI(model="gpt-5.4-mini", timeout=90, max_retries=2),
    tools=[search_docs, lookup_live_docs],
    checkpointer=InMemorySaver(),
    middleware=[
        require_requested_lookup,
        ToolCallLimitMiddleware(run_limit=6),
        ToolCallLimitMiddleware(tool_name="lookup_live_docs", run_limit=3),
    ],
    system_prompt=system_prompt,
)
```

The model can call a tool, read its results, and decide whether it needs more information before answering. The tool limits allow up to six calls per turn, including at most three live lookups. The `@st.cache_resource` decorator on `build_agent()` lets Streamlit reuse the agent when the page reruns.

Live results remain in the conversation without being added to Pinecone. The system prompt asks the model to prefer official sources, but the browsing tool does not enforce a domain restriction.

### Let readers inspect the sources

The **Documentation checked** panel beneath each answer lets readers see which tools ran, whether they returned usable content, and which URLs they retrieved.

`collect_checks()` reads the current turn’s tool responses and extracts their status, source URLs, and any error type. `show_checks()` displays those details in a collapsible panel.

Readers can use the panel alongside the answer’s citations to inspect the evidence. They still need to check whether the returned pages support the instructions.

### Add the chat interface and conversation memory

The final section of `app.py` builds the Streamlit interface. Before accepting questions, it checks for the required API settings and `namespace.txt`. If anything is missing, it explains what to add or run.

Each conversation gets a random `thread_id`. When the reader submits a question, the app passes that question and the thread ID to the agent:

```
result = agent.invoke(
    {"messages": [{"role": "user", "content": question}]},
    config={
        "configurable": {"thread_id": st.session_state.thread_id},
        "recursion_limit": 30,
    },
)
```

LangGraph uses the thread ID to retrieve earlier messages, so `agent.invoke()` only needs the new question. Separately, Streamlit’s session state keeps previous messages and source panels visible when the page reruns.

Clicking **New conversation** creates a new thread ID and clears the visible chat. Conversation memory is temporary: a server restart clears it, and a new browser session starts a new chat.

## Step 5: Run the app and try a conversation

You can now test whether the chatbot functions as intended, at least at the fundamental level: searches the documentation, retrieves additional sources, and remembers your setup.

### Start the app

Save `app.py` (Ctrl + S) and open PyCharm’s **Terminal** in the project folder. On Windows, run:

```
.\.venv\Scripts\python.exe -m streamlit run app.py
```

On macOS or Linux, run:

```
.venv/bin/python -m streamlit run app.py
```

Open the local address shown in the terminal, usually `http://localhost:8501`. Keep the terminal running while you use the chatbot.

### Test the documentation search

The first question gives the chatbot both a task and a constraint:

I have a Python app on an UpCloud server. I want to connect it privately to UpCloud Managed PostgreSQL. Explain what I need to set up.

In the response below, the chatbot explains the private network route and connection setup. The expanded **Documentation checked** panel shows the indexed pages it retrieved, with links to the sources used in the answer.

### Test an external lookup

The next question asks the chatbot to consult documentation outside the indexed collection:

I’m using Psycopg 3. Check the official driver documentation and show me the connection settings, including certificate verification.

Here, the **Documentation checked** panel shows a live lookup through Apify. The chatbot uses the retrieved Psycopg documentation to explain the connection settings, including how to configure certificate verification.

### Test conversation memory

The final question changes where the application runs without repeating the earlier setup:

Actually, during development the app runs on my laptop. What changes?

The response below carries forward the earlier context: UpCloud Managed PostgreSQL, Psycopg 3, and the requirement to keep the database private. It focuses on what changes when the application runs on a laptop, explaining how to reach the private network while preserving certificate verification.

## Step 6: Publish the app with Streamlit

Once the chatbot works locally, you can publish it so others can use it in their browsers. Here's how to set it up:

### Upload the project to GitHub

- Create a [GitHub repository](https://github.com/new) named`upcloud-docs-bot` and initialize it with a README.
- Select **Add file → Upload files** .
- Upload `app.py` ,`ingest.py` ,`requirements.txt` ,`namespace.txt` , and`.gitignore` to the repository’s main folder, alongside the README.
- Commit the files.

Use the `namespace.txt` generated by `ingest.py`. It tells the hosted app which collection to search in Pinecone.

Keep `.env` and `.venv` off GitHub. When uploading through the website, `.gitignore` does not filter the files you select.

### Configure the hosted app

Open [Streamlit Community Cloud](https://share.streamlit.io/), connect your GitHub account, and choose **Create app**.

Enter the repository details:

| Field | Value | 
|---|---|
| Repository | `YOUR_GITHUB_USERNAME/upcloud-docs-bot` | 
| Branch | Your repository’s branch, usually `main` | 
| Main file path | `app.py` | 

### Add the API credentials

Open **Advanced settings** and select **Python 3.12**. Paste the following into **Secrets**, replacing the three credential placeholders:

```
OPENAI_API_KEY = "your_openai_key"
PINECONE_API_KEY = "your_pinecone_key"
APIFY_TOKEN = "your_apify_token"
PINECONE_INDEX_NAME = "upcloud-docs"
```

Keep these settings at the top level, without a section heading. Streamlit makes [root-level secrets available as environment variables](https://docs.streamlit.io/develop/concepts/connections/secrets-management), so the app can read them the same way it reads your local settings.

### Deploy and share

- Save the settings and [deploy the app](https://docs.streamlit.io/deploy/streamlit-community-cloud/deploy-your-app/deploy) . Streamlit will install the dependencies and run`app.py` .
- Open the hosted app and repeat the conversation checks from Step 5.
- Under **Settings → Sharing** , choose who can access the app and copy its URL.

Visitors’ questions use your API accounts, so monitor usage when sharing the app, especially publicly.

The indexed collection reflects your last import. To refresh it:

- Run `ingest.py` again to collect and index the documentation.
- Upload the newly generated `namespace.txt` to GitHub, replacing the previous version.
- Confirm that the hosted chatbot can search the new collection.
- Remove older Pinecone namespaces you no longer need after confirming the new collection works.

## Adapt this chatbot to your documentation

Finding the right guide is only half the work. The harder part is connecting scattered instructions to your setup and keeping that context as the problem evolves.

UpCloud Docs Companion brings those sources into one conversation, lets readers check the guidance, and handles follow-up questions without making them start over or repeat themselves.

To adapt this approach to your documentation, update the starting URLs, crawl filters, coverage checks, and system prompt. Make sure the index dimensions still match the embedding model.

[Create a free Apify account](https://console.apify.com/sign-up) and use the included [$5 in free credit](https://apify.com/pricing) to test your first crawl. Explore thousands of ready-to-run Actors on [Apify Store](https://apify.com/store), including scrapers built for specific websites, to add new sources to your chatbot.

[Contact sales](https://apify.com/contact-sales)
