An overview of the main AI Engineering techniques and patterns used to build, adapt, optimise, and deploy modern AI systems. A developer has compiled an overview of AI engineering techniques, including training from scratch, fine-tuning, quantization, prompt engineering, structured outputs, tool calling, RAG, model routing, and agentic systems. The guide details each technique's advantages, use cases, and requirements to help practitioners choose the right approach for their AI projects. | Technique | Summary | Advantages | When to use | Requirements compute, data | Complexity | |---|---|---|---|---|---| Training from Scratch | Train a new model using your own dataset. | Full control over the model; tailored to your problem. | When no suitable pretrained model exists, or you're building a model for a unique task. | Large datasets and significant compute typically GPUs . | High | Use a Pretrained Model out-of-the-box | Use an existing pretrained model without modifying its weights. | Fastest and simplest approach; no training required; benefits from the capabilities learned during pretraining. | When the model already performs well for the task and additional customisation is unnecessary. | Access to a pretrained model and inference compute. No additional training data required. | Low | Quantization | Reduce numerical precision of model weights e.g. FP16 → INT8 , usually after training. | Smaller models, lower memory usage, faster inference. | Deploying models on limited hardware or reducing inference cost. | No training data; minimal compute. | Low-Medium | Full Fine-tuning | Continue training a pretrained model on your own data updating all the parameters . | Higher accuracy than prompting alone; learns domain-specific behaviour. | You have labelled data and need to adapt a pretrained model. | Labelled dataset and moderate-to-high compute usually GPUs . | High | LoRA / PEFT | Fine-tune only a small subset of model parameters. | Much cheaper and faster than full fine-tuning; small adapter files. | You need to customise a large model with limited resources. | Small labelled dataset and moderate compute. | Medium–High | | Technique | Summary | Advantages | When to use | Requirements compute, data | Complexity | |---|---|---|---|---|---| Prompt Engineering | Design solid instructions and examples to guide model behaviour without changing model weights. | Fast, simple, no retraining required. | The base model already performs reasonably well. | No additional data or compute. | Low | Structured Outputs | Constrain responses to a predefined schema e.g. JSON . | Reliable, machine-readable outputs; easier integration. | Building APIs, automations, or AI applications. | No additional data or compute. | Low | Tool / Function Calling | Allow the model to invoke external tools, APIs, or databases. | Access to real-time data and external capabilities. | The model needs to perform actions or retrieve live information. | APIs, external services, and programming effort. | Medium | RAG Retrieval-Augmented Generation | Retrieve relevant documents and provide them as context to the model. | Uses up-to-date, private, or domain-specific knowledge without retraining. | Answers depend on documents, manuals, PDFs, or internal knowledge bases. | Knowledge base, embeddings, vector database, and indexing pipeline. | Medium–High | Model Routing | Select the most appropriate model for each request. | Reduces cost while maintaining quality. | Different tasks have different quality, latency, or cost requirements. | Access to multiple models and routing logic. | Medium | Agentic Systems / Workflows | Combine planning, memory, and tool use to solve multi-step tasks. | Automates complex workflows and decision-making. | Tasks require multiple reasoning or tool-calling steps. | Multiple tools, orchestration framework, and evaluation. | High |