{"slug": "mastering-langchain-open-source-models-prompt-engineering-part-2", "title": "Mastering LangChain: Open-Source Models & Prompt Engineering (Part 2)", "summary": "A technical tutorial on LangChain demonstrates integrating open-source models via Hugging Face and structuring production prompt templates, including code for connecting to hosted endpoints and local pipelines using models such as Mistral-7B-Instruct-v0.2 and Llama-2-7b-chat-hf.", "body_md": "In Part 1, we established the high-level roadmap of LangChain — exploring the architectural shift toward chat models and breaking down the foundational RAG pipeline.\n\nNow, we dive directly into code implementation. Before assembling end-to-end chains or building multi-step agents, we must master two foundational pillars: **integrating open-source models** and **structuring production prompt templates**.\n\nWhile proprietary APIs (like OpenAI or Anthropic) offer convenience, open-source models provide full data privacy, fine-tuning flexibility, and infrastructure control.\n\nThe largest central hub for discovering, testing, and downloading open-source LLMs is **Hugging Face**. When integrating these models into your LangChain stack, you have two primary options:\n\nFor implementing the below-given codes, you need to get HuggingFace API keys. After you get the API key, copy it and store it in a .env file in the project folder where you will be storing all your LangChain tutorial program files.\n\n**Please refer to the tutorial link on how to get HuggingFace API keys:** [https://www.geeksforgeeks.org/artificial-intelligence/how-to-access-huggingface-api-key/](https://www.geeksforgeeks.org/artificial-intelligence/how-to-access-huggingface-api-key/)\n\nAlso, I have not provided the program outputs for this blog, since I want my readers to actually try the code snippets in their systems.\n\nHere is how you connect both hosted endpoints and local inference pipelines inside LangChain:\n\n``` python\nimport osfrom langchain_huggingface import HuggingFaceEndpoint, HuggingFacePipelinefrom transformers import AutoModelForCausalLM, AutoTokenizer, pipeline# 1. Using Hugging Face Inference API (Serverless / Dedicated Endpoint)os.environ[\"HUGGINGFACEHUB_API_TOKEN\"] = \"your_hf_api_token\"endpoint_llm = HuggingFaceEndpoint(    repo_id=\"mistralai/Mistral-7B-Instruct-v0.2\",    task=\"text-generation\",    max_new_tokens=256,    temperature=0.2)# 2. Running Locally using Transformers Pipelinemodel_id = \"meta-llama/Llama-2-7b-chat-hf\"tokenizer = AutoTokenizer.from_pretrained(model_id)model = AutoModelForCausalLM.from_pretrained(model_id, device_map=\"auto\")pipe = pipeline(    \"text-generation\",    model=model,    tokenizer=tokenizer,    max_new_tokens=256,    temperature=0.3)local_llm = HuggingFacePipeline(pipeline=pipe)\n```\n\nA **prompt** is the input instruction or query provided to a language model to guide its output. In modern conversational architectures, prompts are categorized across distinct roles to establish context, personas, and system boundaries.\n\nNote: We will focus on the LCEL and the formation of chains in langchain in detail in the upcoming blogs. In this blog, I only want you to focus on the prompt part only. You can run the code snippets for now if you wish to do so.\n\nUnlike raw end-user prompts (which are typed directly by users) or system prompts (which define high-level personas and safety guardrails), a **developer prompt** programmatically blends application logic, formatting constraints, and task rules with dynamic user data before sending the payload to the LLM.\n\n**Implementation: Production-Grade Customer Support Triage**\n\nThis implementation demonstrates a developer instruction pipeline that ingests a raw customer support ticket, extracts a concise 2-sentence summary, classifies the urgency into strict enum levels (LOW, MEDIUM & HIGH), and returns a validated Pydantic object\n\n``` python\nimport osfrom langchain_core.prompts import PromptTemplatefrom langchain_core.output_parsers import StrOutputParserfrom langchain_openai import ChatOpenAI# 1. Initialize modelmodel = ChatOpenAI(model=\"gpt-4o-mini\", temperature=0)# 2. Define developer instruction promptdeveloper_prompt = PromptTemplate.from_template(    \"\"\"Summarize the following customer support ticket in exactly 2 sentences.Flag the urgency level strictly as LOW, MEDIUM, or HIGH.Ticket Details:{ticket_text}Output Format:Summary: <summary>Urgency: <urgency>\"\"\")# 3. Create LCEL Chain (Prompt -> Model -> String Parser)chain = developer_prompt | model | StrOutputParser()# 4. Invoke with runtime inputticket_data = \"User cannot log into dashboard after password reset. Multiple attempts failed and account locked.\"response = chain.invoke({\"ticket_text\": ticket_data})print(response)\n```\n\n**Why Developer Prompts Matter in Production**\n\n**2. Multi-Role Conversational Prompt:**\n\nUnlike single-turn base models that treat text as an unstructured continuous string, modern instruction-tuned Chat Models process structured message sequences.\n\nA **multi-role conversational prompt** decomposes an interaction into explicit functional roles (System, User, Assistant), allowing you to steer identity, enforce security guardrails, and dynamically inject historical context without confusing the model.\n\n**Key Roles Breakdown**\n\n**Implementation: Dynamic Multi-Turn Chat with State Management**\n\nIn production, you rarely hardcode past turns. Instead, you use ChatPromptTemplate combined with MessagesPlaceholder and RunnableWithMessageHistory to manage multi-turn conversational state across unique user sessions:\n\n``` python\nfrom langchain_core.prompts import ChatPromptTemplatefrom langchain_core.output_parsers import StrOutputParserfrom langchain_openai import ChatOpenAImodel = ChatOpenAI(model=\"gpt-4o-mini\", temperature=0.2)# Multi-turn chat template with role awarenesschat_prompt = ChatPromptTemplate.from_messages([    (\"system\", \"You are a professional financial advisor for a retail bank. Only discuss bank products, never give speculative advice, and include a non-advice disclaimer.\"),    (\"user\", \"Can you help me understand what a Roth IRA is?\"),    (\"assistant\", \"A Roth IRA is an individual retirement account where you contribute after-tax dollars, allowing tax-free growth and tax-free withdrawals in retirement. (Note: This is general information, not investment advice.)\"),    (\"user\", \"{user_query}\")])chain = chat_prompt | model | StrOutputParser()response = chain.invoke({\"user_query\": \"Does our bank offer automated deposits for a Roth IRA?\"})print(response)\n```\n\n**When to Use Multi-Role Conversational Prompts**\n\n**3. Zero-Shot Prompting**\n\n**Zero-Shot Prompting** is the simplest and most direct prompting technique in Generative AI. In a zero-shot setup, you present a task, query, or instruction to the language model without providing any prior examples or demonstration input-output pairs.\n\nThe model relies entirely on the pre-trained knowledge, linguistic patterns, and semantic understanding it acquired during training to perform the task on the first attempt.\n\n**Implementation: Zero-Shot Sentiment & Intent Classifier**\n\nHere is a complete, production-grade zero-shot implementation using langchain_core and ChatOpenAI with strict enum-based structured outputs to ensure deterministic classification:\n\n``` python\nfrom langchain_core.prompts import PromptTemplatefrom langchain_core.output_parsers import StrOutputParserfrom langchain_openai import ChatOpenAImodel = ChatOpenAI(model=\"gpt-4o-mini\", temperature=0)zero_shot_prompt = PromptTemplate.from_template(    \"\"\"Classify the sentiment of the following customer review strictly as 'Positive', 'Negative', or 'Neutral'.Review: \"{review}\"Sentiment:\"\"\")chain = zero_shot_prompt | model | StrOutputParser()review_input = \"The battery life is decent, but the camera quality disappointed me.\"result = chain.invoke({\"review\": review_input})print(result)\n```\n\n**Key Characteristics**\n\n**4. Few-Shot Prompting**\n\n**Few-Shot Prompting** is a technique where you supply explicit input-output demonstration pairs inside the prompt before presenting the actual query.\n\nWhile zero-shot relies strictly on pre-trained internal patterns, few-shot prompting conditions the model in-context on the exact structure, nuance, formatting constraints, and domain edge cases you expect.\n\n**Implementation: Dynamic Few-Shot Sentiment & Sub-Score Extraction**\n\nLangChain provides dedicated primitives (FewShotChatMessagePromptTemplate and FewShotPromptTemplate) to assemble example pools programmatically:\n\n``` python\nfrom langchain_core.prompts import FewShotPromptTemplate, PromptTemplatefrom langchain_core.output_parsers import StrOutputParserfrom langchain_openai import ChatOpenAImodel = ChatOpenAI(model=\"gpt-4o-mini\", temperature=0)# 1. Provide demonstration examplesexamples = [    {\"review\": \"This product changed my life!\", \"sentiment\": \"Positive\"},    {\"review\": \"Broke after two days, total waste of money.\", \"sentiment\": \"Negative\"}]# 2. Template for each individual exampleexample_prompt = PromptTemplate.from_template(\"Review: \\\"{review}\\\"\\nSentiment: {sentiment}\")# 3. Assemble Few-Shot Templatefew_shot_prompt = FewShotPromptTemplate(    examples=examples,    example_prompt=example_prompt,    prefix=\"Classify sentiment as Positive, Negative, or Neutral.\\n\",    suffix=\"Review: \\\"{review}\\\"\\nSentiment:\",    input_variables=[\"review\"])chain = few_shot_prompt | model | StrOutputParser()result = chain.invoke({\"review\": \"The battery life is decent, but the camera quality disappointed me.\"})print(result)\n```\n\n**5. Chain-of-Thought (CoT) Prompting**\n\nChain-of-thought prompting explicitly asks the model to reason step by step before answering. This can be helpful in solving mathematical problems.\n\n``` python\nfrom langchain_core.prompts import PromptTemplatefrom langchain_core.output_parsers import StrOutputParserfrom langchain_openai import ChatOpenAImodel = ChatOpenAI(model=\"gpt-4o-mini\", temperature=0)cot_prompt = PromptTemplate.from_template(    \"\"\"Solve the following math problem. Think step-by-step and write out your reasoning before stating the final answer.Question: {math_problem}Reasoning:\"\"\")chain = cot_prompt | model | StrOutputParser()problem = \"A store had 120 apples. They sold 45 in the morning and 30 in the afternoon. How many apples are left?\"result = chain.invoke({\"math_problem\": problem})print(result)\n```\n\n**6. Retrieval-Augmented (RAG) Grounding Prompt:**\n\nRestricts model answers strictly to injected reference passages.\n\n``` python\nfrom langchain_core.prompts import PromptTemplatefrom langchain_core.output_parsers import StrOutputParserfrom langchain_openai import ChatOpenAImodel = ChatOpenAI(model=\"gpt-4o-mini\", temperature=0)rag_prompt = PromptTemplate.from_template(    \"\"\"You are a customer support agent. Answer the question using ONLY the provided context below. If the answer cannot be found in the context, output 'Information not found.'Context:{retrieved_context}Question: {user_question}Answer:\"\"\")chain = rag_prompt | model | StrOutputParser()context_data = \"Our refund policy allows returns within 30 days of purchase, with a 15% restocking fee for opened electronics.\"query = \"Can a customer return an opened laptop after 20 days?\"result = chain.invoke({    \"retrieved_context\": context_data,    \"user_question\": query})print(result)\n```\n\n**7. Tree-of-Thought Prompting:**\n\nWhile **Chain-of-Thought (CoT)** forces a model to reason in a straight, linear path, it has a major weakness: if the model makes an incorrect assumption in step 1, that error cascades down the entire chain.\n\n**Tree-of-Thought (ToT)** generalizes over CoT by allowing the model to explore a nonlinearD problem space. It treats reasoning like a search tree where the system generates multiple candidate ideas (branches), evaluates each branch against specific constraints, discards weak paths (pruning), and expands the most promising ones.\n\nImplementation: Building a Tree-of-Thought Pipeline\n\nBecause true Tree-of-Thought requires branching, scoring, and decision-making, it is implemented at the orchestration layer using multi-step chains with structured validation.\n\n``` python\nfrom langchain_core.prompts import ChatPromptTemplatefrom langchain_core.output_parsers import StrOutputParserfrom langchain_openai import ChatOpenAIfrom pydantic import BaseModel, Field# Initialize Modelmodel = ChatOpenAI(model=\"gpt-4o-mini\", temperature=0.7)# --- STEP 1: Branch Generation (Propose Candidate Thoughts) ---generate_branches_prompt = ChatPromptTemplate.from_template(    \"\"\"You are an expert itinerary planner. Generate exactly 3 distinctly different 3-day travel itinerary approaches for the following problem.Problem: {problem}Format output:Approach 1: <summary of focus and day-by-day plan>Approach 2: <summary of focus and day-by-day plan>Approach 3: <summary of focus and day-by-day plan>\"\"\")branch_generator_chain = generate_branches_prompt | model | StrOutputParser()# --- STEP 2: Evaluation & Scoring (Pydantic Structured Schema) ---class ItineraryEvaluation(BaseModel):    chosen_approach_number: int = Field(description=\"The number (1, 2, or 3) of the best approach\")    reasoning: str = Field(description=\"Why this approach best satisfies budget, culture, and feasibility\")    score_out_of_10: float = Field(description=\"Feasibility and quality score out of 10\")evaluator_prompt = ChatPromptTemplate.from_template(    \"\"\"Evaluate the following 3 itinerary options based on time feasibility, cultural balance, and budget efficiency.Candidate Options:{candidates}Select the single most viable candidate branch to expand.\"\"\")evaluator_model = model.with_structured_output(ItineraryEvaluation)evaluator_chain = evaluator_prompt | evaluator_model# --- STEP 3: Branch Expansion (Final Execution on Selected Node) ---expansion_prompt = ChatPromptTemplate.from_template(    \"\"\"Expand the selected travel strategy into a detailed, hourly execution plan including transit details and meal recommendations.Selected Strategy:{selected_plan}\"\"\")expansion_chain = expansion_prompt | model | StrOutputParser()# ==========================================# --- EXECUTION FLOW (Orchestrating ToT) ---# ==========================================problem_statement = \"Plan a 3-day trip to Tokyo balancing historical culture, street food, and budget constraints.\"# 1. Generate Multiple Reasoning Branchesprint(\"Generating candidate thought branches...\")candidates = branch_generator_chain.invoke({\"problem\": problem_statement})print(f\"\\n--- Generated Candidates ---\\n{candidates}\\n\")# 2. Evaluate & Prune Branchesprint(\"Evaluating candidate branches...\")evaluation = evaluator_chain.invoke({\"candidates\": candidates})print(f\"\\n--- Evaluation Verdict ---\\nSelected Branch: Approach {evaluation.chosen_approach_number}\")print(f\"Score: {evaluation.score_out_of_10}/10\\nReason: {evaluation.reasoning}\\n\")# 3. Expand the Winning Branchprint(\"Expanding winning branch into full itinerary...\")final_itinerary = expansion_chain.invoke({\"selected_plan\": candidates})print(f\"\\n--- Final Optimized Output ---\\n{final_itinerary}\")\n```\n\n**When to Use Tree-of-Thought**\n\n**8. Function Calling / Tool-Use (Structured Output Schema)**\n\nBinds JSON schemas to the model so it can return structured tool arguments.\n\n``` python\nfrom langchain_openai import ChatOpenAI# 1. Define Tool Schematools = [    {        \"type\": \"function\",        \"function\": {            \"name\": \"get_stock_price\",            \"description\": \"Fetch current stock price for a ticker symbol.\",            \"parameters\": {                \"type\": \"object\",                \"properties\": {                    \"ticker\": {                        \"type\": \"string\",                        \"description\": \"Stock ticker symbol (e.g., AAPL, MSFT)\"                    }                },                \"required\": [\"ticker\"]            }        }    }]# 2. Bind tool to Chat Modelmodel = ChatOpenAI(model=\"gpt-4o-mini\", temperature=0)model_with_tools = model.bind_tools(tools)# 3. Model decides when to invoke toolresponse = model_with_tools.invoke(\"What is Apple's stock price right now?\")print(response.tool_calls)# Output: [{'name': 'get_stock_price', 'args': {'ticker': 'AAPL'}, 'id': '...'}]\n```\n\nThat is it for today’s blog. Stay tuned as more interesting blogs will come up building upon these concepts.\n\n[Mastering LangChain: Open-Source Models & Prompt Engineering (Part 2)](https://pub.towardsai.net/mastering-langchain-open-source-models-prompt-engineering-part-2-cb4e6e519292) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.", "url": "https://wpnews.pro/news/mastering-langchain-open-source-models-prompt-engineering-part-2", "canonical_source": "https://pub.towardsai.net/mastering-langchain-open-source-models-prompt-engineering-part-2-cb4e6e519292?source=rss----98111c9905da---4", "published_at": "2026-09-08 21:31:02+00:00", "updated_at": "2026-09-08 21:58:48.872497+00:00", "lang": "en", "topics": ["large-language-models", "developer-tools", "ai-tools"], "entities": ["LangChain", "Hugging Face", "Mistral-7B-Instruct-v0.2", "Llama-2-7b-chat-hf", "OpenAI", "Anthropic"], "alternates": {"html": "https://wpnews.pro/news/mastering-langchain-open-source-models-prompt-engineering-part-2", "markdown": "https://wpnews.pro/news/mastering-langchain-open-source-models-prompt-engineering-part-2.md", "text": "https://wpnews.pro/news/mastering-langchain-open-source-models-prompt-engineering-part-2.txt", "jsonld": "https://wpnews.pro/news/mastering-langchain-open-source-models-prompt-engineering-part-2.jsonld"}}