Ollama GPU Scheduling: Running Inference and ComfyUI on One RTX Without OOM A homelab developer documented strategies for sharing a single 16GB RTX 5060 Ti between Ollama LLM inference and ComfyUI Stable Diffusion without running out of VRAM. The writeup catalogs per-model VRAM footprints — Qwen2.5-Coder:14b at roughly 7.5GB, Llama3.1:8b at 4.5GB, SDXL at 6-8GB — and compares time-based versus priority-based GPU scheduling plus static and dynamic VRAM partitioning. It recommends a hybrid of scheduled windows and priority preemption, with smaller 8B models during the day and 14B models in dedicated sessions. Strategies for sharing a single RTX GPU between Ollama LLM inference and ComfyUI Stable Diffusion on the same homelab machine. The VRAM Reality Check: What Actually Fits on a 16GB RTX 5060 Ti Understanding your hardware limits is the first step to successful GPU sharing: Ollama VRAM Consumption Approximate - Qwen2.5-Coder:14b Q4 K M : ~7.5 GB VRAM - DeepSeek-R1:14b Q4 K M : ~7.5 GB VRAM - Llama3.1:8b Q4 K M : ~4.5 GB VRAM - Qwen3:8b Q4 K M : ~4.5 GB VRAM - Nomic-Embed-Text: ~1.2 GB VRAM - System overhead + CUDA context: ~1-2 GB VRAM ComfyUI VRAM Consumption Approximate - Base ComfyUI + VAE: ~1.5-2 GB VRAM - SDXL Base 1024x1024 : ~6-8 GB VRAM - SD 1.5 Base 512x512 : ~3-4 GB VRAM - ControlNet units: ~1-2 GB VRAM each - Hi-Res fix + upscaling: +2-4 GB VRAM - Batch size 1: Linear increase per image - System overhead: ~1-2 GB VRAM The Hard Limits: With 16GB total VRAM: - Ollama 14b + SDXL 1024x1024 = ~7.5GB + ~7GB = ~14.5GB tight but possible - Ollama 14b + SD 1.5 + ControlNet = ~7.5GB + ~4GB + ~2GB = ~13.5GB comfortable - Two 14b models = ~15GB leaves ~1GB for system - not recommended - Ollama 14b + SDXL + Hi-Res + ControlNet = Likely OOM Scheduling Strategies: Time-Based vs Priority-Based Access Two main approaches to GPU sharing, each with different trade-offs: Strategy 1: Time-Based Scheduling Shift Work Allocate explicit time windows to each application: Strategy 2: Priority-Based Access Reservation System Applications request GPU access and get granted based on priority: Hybrid Approach Recommendation: Combine both strategies: - Use time-based scheduling for predictable workloads work hours vs evening - Use priority-based preemption for urgent interactive requests - Allow background tasks to use GPU during idle periods - Implement graceful preemption save state, restore later VRAM Partitioning: Static vs Dynamic Allocation Instead of pure time-sharing, consider splitting the GPU memory: Static Partitioning Fixed Split Divide VRAM upfront and never exceed your allocation: 7GB/9GB Split Ollama Heavy - Ollama: 7GB fits Qwen2.5-Coder:14b or DeepSeek-R1:14b - ComfyUI: 9GB SDXL 1024x1024 + 1 ControlNet + basic upscaling - Best for: Development work where coding assistance is primary 5GB/11GB Split Balanced - Ollama: 5GB fits Llama3.1:8b or Qwen3:8b - ComfyUI: 11GB SDXL + multiple ControlNets + Hi-Res fix - Best for: Mixed usage where both get reasonable resources 3GB/13GB Split ComfyUI Heavy - Ollama: 3GB very small models or embedding-only - ComfyUI: 13GB SDXL batch processing, video generation, extensive ControlNet - Best for: Art generation focus with occasional LLM queries Static Partitioning Limitations: While simple, this approach has drawbacks: - Wasted resources when one application is idle - No ability to handle bursts beyond allocation - Requires restarting applications to change partition sizes - Complex to implement correctly with CUDA/Vulkan Dynamic Partitioning Preferred Approach Allow applications to use available VRAM, but with limits and priorities: Practical Dynamic Strategy: For most homelab users: - Run Ollama with smaller models during the day 8b parameters - Switch to larger models 14b during dedicated LLM sessions - Use ComfyUI with moderate settings most of the time - Save extreme ComfyUI settings for dedicated creative sessions - Use VRAM monitoring to avoid surprises Model Quantization: Your Best Friend for VRAM Efficiency Choosing the right quantization level dramatically affects what fits: Quantization Impact on 14b Models | Quantization | VRAM Usage | Quality Impact | Speed Impact | | Q8 0 | ~10.5 GB | Minimal | None | | Q6 K | ~9.0 GB | Very minor | None | | Q5 K M | ~7.8 GB | Minor | None | | Q4 K M | ~7.5 GB | Noticeable but acceptable | None | | Q3 K M | ~6.2 GB | Moderate | None | | Q2 K | ~5.0 GB | Significant | None | Practical Quantization Recommendations - For coding assistance: Q4 K M or Q5 K M good balance - For mathematical/logical tasks: Q5 K M or Q6 K better reasoning - For casual chat: Q3 K M or Q2 K acceptable quality, saves VRAM - For embeddings: Always use quantized versions tiny VRAM footprint - Avoid: Floating point F16 unless you have 24GB+ VRAM ComfyUI Optimization: Getting More from Limited VRAM Stable Diffusion has many knobs to tune for VRAM efficiency: Essential ComfyUI VRAM Savers: - Use SD 1.5 instead of SDXL when possible: ~50% VRAM reduction - Enable attention slicing: trades compute for VRAM often worth it - Use VAE tiling: processes image in chunks to reduce peak VRAM - Limit ControlNet units: each additional unit adds ~1-2GB VRAM - Use lower batch sizes: batch=1 is VRAM efficient - Enable model offloading: move unused components to RAM - Use xformers: more memory-efficient attention implementation - Prefer FP8 over FP16 when available: halves VRAM for certain operations Resolution and Performance Trade-offs: VRAM by Resolution SD 1.5 | Resolution | Base VRAM | With Hi-Res Fix | Typical Use Case | | 512x512 | ~3.5 GB | ~5.5 GB | Quick iterations, concepts | | 768x768 | ~5.0 GB | ~8.0 GB | Balanced quality/speed | | 1024x1024 | ~7.5 GB | ~12.0 GB | High quality generations | | 1152x896 | ~8.5 GB | ~13.5 GB | Widescreen compositions | VRAM by Resolution SDXL | Resolution | Base VRAM | With Hi-Res Fix | Typical Use Case | | 768x768 | ~6.0 GB | ~9.0 GB | Moderate quality | | 1024x1024 | ~8.0 GB | ~12.5 GB | Standard quality | | 1152x896 | ~9.0 GB | ~14.0 GB | Widescreen needs optimization | | 1280x720 | ~9.5 GB | ~15.0 GB | Landscape tight fit | Queue Management: Preventing Resource Exhaustion Even with good scheduling, prevent overload with proper queueing: Ollama Request Queuing: Prevent overwhelming the LLM with concurrent requests: ComfyUI Job Queuing: Manage image generation requests to prevent VRAM spikes: Monitoring and Alerting: Knowing When Things Go Wrong Essential visibility into your shared GPU setup: Key Metrics to Monitor: - VRAM Usage: Total, used, free with alerts at 80%, 90%, 95% - GPU Utilization: Percentage of time GPU is actively computing - Temperature: GPU and memory temperatures throttling risks - Power Draw: Watts consumed PSU capacity planning - Application Response Times: Ollama latency, ComfyUI generation time - Queue Depths: Number of pending requests for each service - Error Rates: Failed generations, OOM occurrences, timeouts Practical Daily Workflow: Putting It All Together How to structure your day for optimal shared GPU usage: Developer-Focused Day 6:00 AM - 12:00 PM: Coding Focus - Ollama: Qwen2.5-Coder:14b-q4 K M 7.5GB VRAM - ComfyUI: Limited to SD 1.5 512x512, batch=1 4GB VRAM - Usage: Quick concept illustrations, diagram generation 12:00 PM - 1:00 PM: Lunch Break - Either app can burst to full VRAM if needed - Good time for experimental generations or large model tests 1:00 PM - 6:00 PM: Continued Development - Same as morning session - Ollama might handle more complex reasoning tasks Artist-Focused Day 6:00 AM - 12:00 PM: Creation Focus - ComfyUI: SDXL 1024x1024 + ControlNet + Hi-Res 10GB VRAM - Ollama: Llama3.1:8b-q4 K M for prompt assistance 4.5GB VRAM - Usage: Generating variations, refining prompts, quick idea exploration 12:00 PM - 1:00 PM: Lunch Break - Resource sharing/experimentation time 1:00 PM - 6:00 PM: Continued Creation - Same as morning session - Might batch process multiple variations Balanced/Hybrid Day 6:00 AM - 10:00 AM: Mixed Light Usage - Ollama: Qwen2.5-Coder:8b for light assistance - ComfyUI: SD 1.5 768x768 for occasional illustrations 10:00 AM - 12:00 PM: Ollama Heavy - Switch to Qwen2.5-Coder:14b for deep work - ComfyUI minimal or idle 12:00 PM - 2:00 PM: Lunch/Experimentation - Try larger models or experimental settings 2:00 PM - 6:00 PM: ComfyUI Focused - ComfyUI: SDXL + experiments - Ollama: 8b model for light background tasks Troubleshooting Common Issues Solutions to problems you'll encounter: Frequent Out-of-Memory Errors - Solution: Implement graceful degradation - Implementation: Slow Performance When Switching - Cause: Model reloading, VRAM fragmentation - Solutions: Unresponsive Applications - Cause: Deadlock, infinite waits, resource starvation - Solutions: VRAM Fragmentation Issues - Cause: Repeated allocation/deallocation patterns - Solutions: Conclusion: Making One GPU Work for Both Worlds Sharing a single RTX GPU between Ollama and ComfyUI is entirely practical with the right approach: Key Principles to Remember: - Know your limits: Understand exactly what fits in your 16GB VRAM - Quantization is your friend: The difference between Q4 and Q5 can make or break your setup - Schedule intentionally: Don't let both applications fight for resources randomly - Monitor relentlessly: Visibility prevents surprises and enables optimization - Graceful degradation hard failure: Build systems that slow down rather than crash - Experiment to find your sweet spot: Your workload patterns are unique Recommended Starting Configuration: - Ollama: Qwen2.5-Coder:14b-q4 K M ~7.5GB or Llama3.1:8b-q4 K M ~4.5GB - ComfyUI: SD 1.5 with attention slicing and VAE tiling enabled - Scheduling: Time-based with priority preemption for urgent requests - Monitoring: Basic VRAM usage alerts + application health checks - Workflow: Batch similar tasks together to minimize context switching costs Final Thought: The goal isn't to run both systems at maximum capacity simultaneously — that's unrealistic on consumer hardware. The goal is to have both systems available when you need them, with predictable performance and minimal frustration. With thoughtful scheduling, appropriate model choices, and basic monitoring, your RTX 5060 Ti can serve as a capable dual-purpose AI workstation for both development and creative work. Originally published on ayraix.com https://ayraix.com/signal/community/ollama-gpu-scheduling-rtx-comfyui-2026/ , practical AI for enterprise builders.