Automating Knowledge Graph Population: Extracting Entities and Triples from Unstructured Text with an LLM A tutorial published on Machine Learning Mastery details a pipeline that uses a local Llama 3.2 model served through Ollama to extract SPOC (Subject-Predicate-Object-Context) quads from unstructured Wikipedia text and load them into the Quadstore knowledge graph database for Graph-RAG retrieval. The workflow runs in Google Colab or a local Python IDE, requires installing the wikipedia and requests libraries, and relies on few-shot prompting with Llama 3.2's JSON output mode to convert narrative text into structured relationships. In this article, you will learn how to automatically extract structured knowledge from raw text and populate a knowledge graph with SPOC quads using a local LLM via Ollama. Topics we will cover include: - How to set up Ollama with the Llama 3.2 model to run a free, local LLM for structured data extraction. - How to design a robust extraction pipeline that converts unstructured Wikipedia text into SPOC Subject-Predicate-Object-Context quads using few-shot prompting and JSON output mode. - How to load the extracted quads into a QuadStore knowledge graph, ready for use in a Graph-RAG retrieval pipeline. Introduction The recent article on Building a Deterministic 3-Tiered Graph-RAG System https://machinelearningmastery.com/beyond-vector-search-building-a-deterministic-3-tiered-graph-rag-system/ shows how a hierarchical, graph-based architecture can tackle the issue of hallucinations in standard vector information retrieval. That article leveraged Quadstore https://github.com/mmmayo13/quadstore , a lightweight knowledge graph database implemented in Python, to teach LLMs to respect ground-truth facts, thereby ensuring factual accuracy and deterministic retrieval conflict resolution in applications like RAG systems. A critical question remains, though: where does the factual graph knowledge come from? This article helps close the loop, showing a free, fully automated approach to extract entities and build SPOC quads Subject-Predicate-Object-Context from raw text such as Wikipedia pages. We will do this with the help of a local, free LLM from Ollama. Once these quads are built, we will illustrate how to directly populate the Quadstore. Prerequisites and Setup The workflow shown in this article is designed to run seamlessly both in a Google Colab notebook and in your local Python IDE. If you choose the latter, you will need to manually install Ollama on your computer first, along with pulling the Llama 3.2 model locally. In Google Colab, you can set up Ollama and get Llama 3.2 for your open session using these commands: apt-get update -qq && apt-get install -y -qq zstd curl -fsSL https://ollama.com/install.sh | sh 12 apt-get update -qq && apt-get install -y -qq zstd curl -fsSL https://ollama.com/install.sh | sh Either way, you will need to install these two libraries as well: pip install wikipedia requests 1 pip install wikipedia requests Llama 3.2 is a lightweight, free model. Its API is configured to strictly operate in JSON input/output mode, a mandatory standard for reliable data extraction. Using subprocess , we can start the Ollama server as a background process and pull our target model: python import subprocess import time 1. Starting the Ollama server in the background print "Starting Ollama server..." process = subprocess.Popen "ollama", "serve" , stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL time.sleep 3 Give the server a moment to start 2. Pulling the Llama 3.2 model this may take a minute or two on Colab print "Pulling Llama 3.2..." subprocess.run "ollama", "pull", "llama3.2" print "Model ready " 123456789101112 import subprocessimport time 1. Starting the Ollama server in the backgroundprint "Starting Ollama server..." process = subprocess.Popen "ollama", "serve" , stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL time.sleep 3 Give the server a moment to start 2. Pulling the Llama 3.2 model this may take a minute or two on Colab print "Pulling Llama 3.2..." subprocess.run "ollama", "pull", "llama3.2" print "Model ready " Automated Knowledge Graph Population Let’s look at the process of constructing our knowledge graph from a source of raw text. We will convert narrative text into strictly modeled relationships that extend classical RDF triples of the form Subject, Predicate, Object by adding a fourth dimension: the context. This is useful for tracking where a fact comes from and whether it is true or not. A triple like "LeBron James", "plays for", "Lakers" thus becomes "LeBron James", "plays for", "Lakers", "NBA 2023 Roster" . First, we will create a small, simulated QuadStore engine that mimics the framework used in the related article this one follows up on. If you are working in a notebook, run this code in a separate cell to generate a quadstore.py file in your workspace on the fly: python %%writefile quadstore.py class QuadStore: def init self : Using a simple list to store the facts for our lightweight implementation self.quads = def add self, subject, predicate, obj, context : """Adds a new SPOC quad to the knowledge graph.""" quad = subject, predicate, obj, context if quad not in self.quads: self.quads.append quad def query self, subject=None, predicate=None, obj=None, context=None : """Queries the graph. Returns a list of quads that match the provided criteria.""" results = for q sub, q pred, q obj, q ctx in self.quads: if subject is None or subject == q sub and \ predicate is None or predicate == q pred and \ obj is None or obj == q obj and \ context is None or context == q ctx : results.append q sub, q pred, q obj, q ctx return results 1234567891011121314151617181920212223 %%writefile quadstore.py class QuadStore: def init self : Using a simple list to store the facts for our lightweight implementation self.quads = def add self, subject, predicate, obj, context : """Adds a new SPOC quad to the knowledge graph.""" quad = subject, predicate, obj, context if quad not in self.quads: self.quads.append quad def query self, subject=None, predicate=None, obj=None, context=None : """Queries the graph. Returns a list of quads that match the provided criteria.""" results = for q sub, q pred, q obj, q ctx in self.quads: if subject is None or subject == q sub and \ predicate is None or predicate == q pred and \ obj is None or obj == q obj and \ context is None or context == q ctx : results.append q sub, q pred, q obj, q ctx return results A file containing exactly the above code will be created, and we will refer to it later on just like any other Python module, to demonstrate how to load our created knowledge graph into our mock Graph-RAG system. Back to the main process: we will now pull some raw text from Wikipedia using the namesake API. The auto suggest=False option ensures we correctly fetch the right article name from Wikipedia without automated corrections that may cause a crash. python import wikipedia print "Fetching Wikipedia summary..." Disabling auto suggest to prevent the library from renaming "Turing" to "tuning" wiki page = wikipedia.page "Alan Turing", auto suggest=False text content = wiki page.summary We'll just take the first two paragraphs of the Wikipedia page to keep extraction fast paragraphs = text content.split '\n' :2 short text = " ".join paragraphs print f"Extracted {len short text } characters of text ready for processing." 123456789101112 import wikipedia print "Fetching Wikipedia summary..." Disabling auto suggest to prevent the library from renaming "Turing" to "tuning"wiki page = wikipedia.page "Alan Turing", auto suggest=False text content = wiki page.summary We'll just take the first two paragraphs of the Wikipedia page to keep extraction fastparagraphs = text content.split '\n' :2 short text = " ".join paragraphs print f"Extracted {len short text } characters of text ready for processing." Output: Fetching Wikipedia summary... Extracted 1244 characters of text ready for processing. 12 Fetching Wikipedia summary...Extracted 1244 characters of text ready for processing. Now comes the core of the entire workflow: the robust extraction engine, modeled by the following function that: - Works with the target LLM’s formatting engine to extract structured data in the form of quads. To do this, we use few-shot examples as part of the prompt sent to the LLM. - Post-processes the LLM output to extract a list of facts and build a list of quads accordingly. python import json import requests def extract spoc quads final text, context label, model="llama3.2" : To abide by Llama3.2's output mode, we ask for a JSON object with a "facts" key prompt = f""" You are an expert data extraction algorithm. Extract atomic facts from the text. You must output a valid JSON object containing a single key called "facts". The value of "facts" must be an array of objects. Example output format: {{ "facts": {{"subject": "LeBron James", "predicate": "plays for", "object": "Lakers"}}, {{"subject": "Lakers", "predicate": "based in", "object": "Los Angeles"}} }} Text to process: {text} """ payload = { "model": model, "prompt": prompt, "format": "json", "stream": False, "temperature": 0.0 } try: response = requests.post 'http://localhost:11434/api/generate', json=payload response.raise for status raw llm text = response.json 'response' parsed json = json.loads raw llm text Looking specifically for the "facts" array triples = parsed json.get "facts", Fallback: If the LLM still used a different key, grab the first list we find if not triples and isinstance parsed json, dict : for key, value in parsed json.items : if isinstance value, list : triples = value break quads = for t in triples: if not isinstance t, dict : continue Converting keys to lowercase to catch "Subject" vs "subject" normalized t = {str k .lower .strip : str v .strip for k, v in t.items } if all k in normalized t for k in 'subject', 'predicate', 'object' : quads.append { "subject": normalized t 'subject' , "predicate": normalized t 'predicate' , "object": normalized t 'object' , "context": context label } return quads except Exception as e: print f"Extraction failed: {e}" return 123456789101112131415161718192021222324252627282930313233343536373839404142434445464748495051525354555657585960616263646566 import jsonimport requests def extract spoc quads final text, context label, model="llama3.2" : To abide by Llama3.2's output mode, we ask for a JSON object with a "facts" key prompt = f""" You are an expert data extraction algorithm. Extract atomic facts from the text. You must output a valid JSON object containing a single key called "facts". The value of "facts" must be an array of objects. Example output format: {{ "facts": {{"subject": "LeBron James", "predicate": "plays for", "object": "Lakers"}}, {{"subject": "Lakers", "predicate": "based in", "object": "Los Angeles"}} }} Text to process: {text} """ payload = { "model": model, "prompt": prompt, "format": "json", "stream": False, "temperature": 0.0 } try: response = requests.post 'http://localhost:11434/api/generate', json=payload response.raise for status raw llm text = response.json 'response' parsed json = json.loads raw llm text Looking specifically for the "facts" array triples = parsed json.get "facts", Fallback: If the LLM still used a different key, grab the first list we find if not triples and isinstance parsed json, dict : for key, value in parsed json.items : if isinstance value, list : triples = value break quads = for t in triples: if not isinstance t, dict : continue Converting keys to lowercase to catch "Subject" vs "subject" normalized t = {str k .lower .strip : str v .strip for k, v in t.items } if all k in normalized t for k in 'subject', 'predicate', 'object' : quads.append { "subject": normalized t 'subject' , "predicate": normalized t 'predicate' , "object": normalized t 'object' , "context": context label } return quads except Exception as e: print f"Extraction failed: {e}" return All that remains is running the pipeline to extract, view, and make use of our newly created quads, which will constitute our knowledge graph. print "Beginning extraction this takes a few seconds on a Colab T4 GPU ...\n" extracted quads = extract spoc quads final text=short text, context label="Wikipedia Alan Turing" Viewing the results for quad in extracted quads: print f"S: {quad 'subject' :<20} | P: {quad 'predicate' :<15} | O: {quad 'object' :<25} | C: {quad 'context' }" 12345678910 print "Beginning extraction this takes a few seconds on a Colab T4 GPU ...\n" extracted quads = extract spoc quads final text=short text, context label="Wikipedia Alan Turing" Viewing the resultsfor quad in extracted quads: print f"S: {quad 'subject' :<20} | P: {quad 'predicate' :<15} | O: {quad 'object' :<25} | C: {quad 'context' }" Results: Beginning extraction this takes a few seconds on a Colab T4 GPU ... S: Alan Mathison Turing | P: was | O: an English mathematician, computer scientist, logician, cryptanalyst, philosopher and theoretical biologist | C: Wikipedia Alan Turing S: Alan Mathison Turing | P: was born | O: in London | C: Wikipedia Alan Turing S: Alan Mathison Turing | P: was raised | O: in southern England | C: Wikipedia Alan Turing S: Alan Mathison Turing | P: graduated from | O: King's College, Cambridge | C: Wikipedia Alan Turing S: Alan Mathison Turing | P: earned | O: a doctorate degree from Princeton University | C: Wikipedia Alan Turing S: Alan Mathison Turing | P: worked for | O: the Government Code and Cypher School at Bletchley Park | C: Wikipedia Alan Turing S: Alan Mathison Turing | P: led | O: Hut 8, the section responsible for German naval cryptanalysis | C: Wikipedia Alan Turing S: Alan Mathison Turing | P: devised techniques for | O: speeding the breaking of German ciphers | C: Wikipedia Alan Turing S: Alan Mathison Turing | P: played a crucial role in | O: cracking intercepted messages that enabled the Allies to defeat the Axis powers | C: Wikipedia Alan Turing S: Alan Mathison Turing | P: is | O: widely considered to be the father of theoretical computer science | C: Wikipedia Alan Turing S: Alan Mathison Turing | P: was | O: influential in the development of theoretical computer science | C: Wikipedia Alan Turing 12345678910111213 Beginning extraction this takes a few seconds on a Colab T4 GPU ... S: Alan Mathison Turing | P: was | O: an English mathematician, computer scientist, logician, cryptanalyst, philosopher and theoretical biologist | C: Wikipedia Alan TuringS: Alan Mathison Turing | P: was born | O: in London | C: Wikipedia Alan TuringS: Alan Mathison Turing | P: was raised | O: in southern England | C: Wikipedia Alan TuringS: Alan Mathison Turing | P: graduated from | O: King's College, Cambridge | C: Wikipedia Alan TuringS: Alan Mathison Turing | P: earned | O: a doctorate degree from Princeton University | C: Wikipedia Alan TuringS: Alan Mathison Turing | P: worked for | O: the Government Code and Cypher School at Bletchley Park | C: Wikipedia Alan TuringS: Alan Mathison Turing | P: led | O: Hut 8, the section responsible for German naval cryptanalysis | C: Wikipedia Alan TuringS: Alan Mathison Turing | P: devised techniques for | O: speeding the breaking of German ciphers | C: Wikipedia Alan TuringS: Alan Mathison Turing | P: played a crucial role in | O: cracking intercepted messages that enabled the Allies to defeat the Axis powers | C: Wikipedia Alan TuringS: Alan Mathison Turing | P: is | O: widely considered to be the father of theoretical computer science | C: Wikipedia Alan TuringS: Alan Mathison Turing | P: was | O: influential in the development of theoretical computer science | C: Wikipedia Alan Turing From just two paragraphs of Alan Turing’s Wikipedia article, we extracted around 11 facts, structured as quads. Note that the exact number may vary slightly due to the non-deterministic behavior of the LLM. We wrap up by seeing how to add these facts into the QuadStore object: python from quadstore import QuadStore facts qs = QuadStore for quad in extracted quads: facts qs.add quad "subject" , quad "predicate" , quad "object" , quad "context" print f"Successfully loaded {len extracted quads } automated facts into the Graph RAG system " 12345678910111213 from quadstore import QuadStore facts qs = QuadStore for quad in extracted quads: facts qs.add quad "subject" , quad "predicate" , quad "object" , quad "context" print f"Successfully loaded {len extracted quads } automated facts into the Graph RAG system " Output: Successfully loaded 11 automated facts into the Graph RAG system 1 Successfully loaded 11 automated facts into the Graph RAG system Conclusion This article closed the loop on our deterministic 3-tiered Graph-RAG architecture https://machinelearningmastery.com/beyond-vector-search-building-a-deterministic-3-tiered-graph-rag-system/ by showing how to build a knowledge graph consisting of facts extracted directly from unstructured text in the form of quads — all from scratch. You can now integrate this knowledge into your retrieval pipeline to help eliminate issues like LLM hallucinations.