# Ollama GPU Scheduling: Running Inference and ComfyUI on One RTX Without OOM

> Source: <https://dev.to/ayraix/ollama-gpu-scheduling-running-inference-and-comfyui-on-one-rtx-without-oom-1j75>
> Published: 2026-10-11 21:00:00+00:00

*Strategies for sharing a single RTX GPU between Ollama LLM inference and ComfyUI Stable Diffusion on the same homelab machine.*

## 
  
  
  The VRAM Reality Check: What Actually Fits on a 16GB RTX 5060 Ti

Understanding your hardware limits is the first step to successful GPU sharing:

### 
  
  
  Ollama VRAM Consumption (Approximate)

- 
**Qwen2.5-Coder:14b (Q4_K_M):** ~7.5 GB VRAM
- 
**DeepSeek-R1:14b (Q4_K_M):** ~7.5 GB VRAM
- 
**Llama3.1:8b (Q4_K_M):** ~4.5 GB VRAM
- 
**Qwen3:8b (Q4_K_M):** ~4.5 GB VRAM
- 
**Nomic-Embed-Text:** ~1.2 GB VRAM
- 
**System overhead + CUDA context:** ~1-2 GB VRAM

### 
  
  
  ComfyUI VRAM Consumption (Approximate)

- 
**Base ComfyUI + VAE:** ~1.5-2 GB VRAM
- 
**SDXL Base (1024x1024):** ~6-8 GB VRAM
- 
**SD 1.5 Base (512x512):** ~3-4 GB VRAM
- 
**ControlNet units:** ~1-2 GB VRAM each
- 
**Hi-Res fix + upscaling:** +2-4 GB VRAM
- 
**Batch size >1:** Linear increase per image
- 
**System overhead:** ~1-2 GB VRAM

### 
  
  
  The Hard Limits:

With 16GB total VRAM:

- Ollama 14b + SDXL 1024x1024 = ~7.5GB + ~7GB = ~14.5GB (tight but possible)
- Ollama 14b + SD 1.5 + ControlNet = ~7.5GB + ~4GB + ~2GB = ~13.5GB (comfortable)
- Two 14b models = ~15GB (leaves ~1GB for system - not recommended)
- Ollama 14b + SDXL + Hi-Res + ControlNet = Likely OOM

## 
  
  
  Scheduling Strategies: Time-Based vs Priority-Based Access

Two main approaches to GPU sharing, each with different trade-offs:

### 
  
  
  Strategy 1: Time-Based Scheduling (Shift Work)

Allocate explicit time windows to each application:

### 
  
  
  Strategy 2: Priority-Based Access (Reservation System)

Applications request GPU access and get granted based on priority:

### 
  
  
  Hybrid Approach Recommendation:

Combine both strategies:

- Use time-based scheduling for predictable workloads (work hours vs evening)
- Use priority-based preemption for urgent interactive requests
- Allow background tasks to use GPU during idle periods
- Implement graceful preemption (save state, restore later)

## 
  
  
  VRAM Partitioning: Static vs Dynamic Allocation

Instead of pure time-sharing, consider splitting the GPU memory:

### 
  
  
  Static Partitioning (Fixed Split)

Divide VRAM upfront and never exceed your allocation:

#### 
  
  
  7GB/9GB Split (Ollama Heavy)

- Ollama: 7GB (fits Qwen2.5-Coder:14b or DeepSeek-R1:14b)
- ComfyUI: 9GB (SDXL 1024x1024 + 1 ControlNet + basic upscaling)
- Best for: Development work where coding assistance is primary

#### 
  
  
  5GB/11GB Split (Balanced)

- Ollama: 5GB (fits Llama3.1:8b or Qwen3:8b)
- ComfyUI: 11GB (SDXL + multiple ControlNets + Hi-Res fix)
- Best for: Mixed usage where both get reasonable resources

#### 
  
  
  3GB/13GB Split (ComfyUI Heavy)

- Ollama: 3GB (very small models or embedding-only)
- ComfyUI: 13GB (SDXL batch processing, video generation, extensive ControlNet)
- Best for: Art generation focus with occasional LLM queries

### 
  
  
  Static Partitioning Limitations:

While simple, this approach has drawbacks:

- Wasted resources when one application is idle
- No ability to handle bursts beyond allocation
- Requires restarting applications to change partition sizes
- Complex to implement correctly with CUDA/Vulkan

### 
  
  
  Dynamic Partitioning (Preferred Approach)

Allow applications to use available VRAM, but with limits and priorities:

### 
  
  
  Practical Dynamic Strategy:

For most homelab users:

- Run Ollama with smaller models during the day (8b parameters)
- Switch to larger models (14b) during dedicated LLM sessions
- Use ComfyUI with moderate settings most of the time
- Save extreme ComfyUI settings for dedicated creative sessions
- Use VRAM monitoring to avoid surprises

## 
  
  
  Model Quantization: Your Best Friend for VRAM Efficiency

Choosing the right quantization level dramatically affects what fits:

### 
  
  
  Quantization Impact on 14b Models

| Quantization | VRAM Usage | Quality Impact | Speed Impact | 
| Q8_0 | ~10.5 GB | Minimal | None | 
| Q6_K | ~9.0 GB | Very minor | None | 
| Q5_K_M | ~7.8 GB | Minor | None | 
| Q4_K_M | ~7.5 GB | Noticeable but acceptable | None | 
| Q3_K_M | ~6.2 GB | Moderate | None | 
| Q2_K | ~5.0 GB | Significant | None | 

### 
  
  
  Practical Quantization Recommendations

- 
**For coding assistance:** Q4_K_M or Q5_K_M (good balance)
- 
**For mathematical/logical tasks:** Q5_K_M or Q6_K (better reasoning)
- 
**For casual chat:** Q3_K_M or Q2_K (acceptable quality, saves VRAM)
- 
**For embeddings:** Always use quantized versions (tiny VRAM footprint)
- 
**Avoid:** Floating point (F16) unless you have 24GB+ VRAM

## 
  
  
  ComfyUI Optimization: Getting More from Limited VRAM

Stable Diffusion has many knobs to tune for VRAM efficiency:

### 
  
  
  Essential ComfyUI VRAM Savers:

- 
**Use SD 1.5 instead of SDXL when possible:** ~50% VRAM reduction
- 
**Enable attention slicing:** trades compute for VRAM (often worth it)
- 
**Use VAE tiling:** processes image in chunks to reduce peak VRAM
- 
**Limit ControlNet units:** each additional unit adds ~1-2GB VRAM
- 
**Use lower batch sizes:** batch=1 is VRAM efficient
- 
**Enable model offloading:** move unused components to RAM
- 
**Use xformers:** more memory-efficient attention implementation
- 
**Prefer FP8 over FP16 when available:** halves VRAM for certain operations

### 
  
  
  Resolution and Performance Trade-offs:

#### 
  
  
  VRAM by Resolution (SD 1.5)

| Resolution | Base VRAM | With Hi-Res Fix | Typical Use Case | 
| 512x512 | ~3.5 GB | ~5.5 GB | Quick iterations, concepts | 
| 768x768 | ~5.0 GB | ~8.0 GB | Balanced quality/speed | 
| 1024x1024 | ~7.5 GB | ~12.0 GB | High quality generations | 
| 1152x896 | ~8.5 GB | ~13.5 GB | Widescreen compositions | 

#### 
  
  
  VRAM by Resolution (SDXL)

| Resolution | Base VRAM | With Hi-Res Fix | Typical Use Case | 
| 768x768 | ~6.0 GB | ~9.0 GB | Moderate quality | 
| 1024x1024 | ~8.0 GB | ~12.5 GB | Standard quality | 
| 1152x896 | ~9.0 GB | ~14.0 GB | Widescreen (needs optimization) | 
| 1280x720 | ~9.5 GB | ~15.0 GB | Landscape (tight fit) | 

## 
  
  
  Queue Management: Preventing Resource Exhaustion

Even with good scheduling, prevent overload with proper queueing:

### 
  
  
  Ollama Request Queuing:

Prevent overwhelming the LLM with concurrent requests:

### 
  
  
  ComfyUI Job Queuing:

Manage image generation requests to prevent VRAM spikes:

## 
  
  
  Monitoring and Alerting: Knowing When Things Go Wrong

Essential visibility into your shared GPU setup:

### 
  
  
  Key Metrics to Monitor:

- 
**VRAM Usage:** Total, used, free (with alerts at 80%, 90%, 95%)
- 
**GPU Utilization:** Percentage of time GPU is actively computing
- 
**Temperature:** GPU and memory temperatures (throttling risks)
- 
**Power Draw:** Watts consumed (PSU capacity planning)
- 
**Application Response Times:** Ollama latency, ComfyUI generation time
- 
**Queue Depths:** Number of pending requests for each service
- 
**Error Rates:** Failed generations, OOM occurrences, timeouts

## 
  
  
  Practical Daily Workflow: Putting It All Together

How to structure your day for optimal shared GPU usage:

### 
  
  
  Developer-Focused Day

**6:00 AM - 12:00 PM:** Coding Focus

- Ollama: Qwen2.5-Coder:14b-q4_K_M (7.5GB VRAM)
- ComfyUI: Limited to SD 1.5 512x512, batch=1 (4GB VRAM)
- Usage: Quick concept illustrations, diagram generation

**12:00 PM - 1:00 PM:** Lunch Break

- Either app can burst to full VRAM if needed
- Good time for experimental generations or large model tests

**1:00 PM - 6:00 PM:** Continued Development

- Same as morning session
- Ollama might handle more complex reasoning tasks

### 
  
  
  Artist-Focused Day

**6:00 AM - 12:00 PM:** Creation Focus

- ComfyUI: SDXL 1024x1024 + ControlNet + Hi-Res (10GB VRAM)
- Ollama: Llama3.1:8b-q4_K_M for prompt assistance (4.5GB VRAM)
- Usage: Generating variations, refining prompts, quick idea exploration

**12:00 PM - 1:00 PM:** Lunch Break

- Resource sharing/experimentation time

**1:00 PM - 6:00 PM:** Continued Creation

- Same as morning session
- Might batch process multiple variations

### 
  
  
  Balanced/Hybrid Day

**6:00 AM - 10:00 AM:** Mixed Light Usage

- Ollama: Qwen2.5-Coder:8b for light assistance
- ComfyUI: SD 1.5 768x768 for occasional illustrations

**10:00 AM - 12:00 PM:** Ollama Heavy

- Switch to Qwen2.5-Coder:14b for deep work
- ComfyUI minimal or idle

**12:00 PM - 2:00 PM:** Lunch/Experimentation

- Try larger models or experimental settings

**2:00 PM - 6:00 PM:** ComfyUI Focused

- ComfyUI: SDXL + experiments
- Ollama: 8b model for light background tasks

## 
  
  
  Troubleshooting Common Issues

Solutions to problems you'll encounter:

### 
  
  
  Frequent Out-of-Memory Errors

- 
**Solution:** Implement graceful degradation
- 
**Implementation:**

### 
  
  
  Slow Performance When Switching

- 
**Cause:** Model reloading, VRAM fragmentation
- 
**Solutions:**

### 
  
  
  Unresponsive Applications

- 
**Cause:** Deadlock, infinite waits, resource starvation
- 
**Solutions:**

### 
  
  
  VRAM Fragmentation Issues

- 
**Cause:** Repeated allocation/deallocation patterns
- 
**Solutions:**

## 
  
  
  Conclusion: Making One GPU Work for Both Worlds

Sharing a single RTX GPU between Ollama and ComfyUI is entirely practical with the right approach:

### 
  
  
  Key Principles to Remember:

- 
**Know your limits:** Understand exactly what fits in your 16GB VRAM
- 
**Quantization is your friend:** The difference between Q4 and Q5 can make or break your setup
- 
**Schedule intentionally:** Don't let both applications fight for resources randomly
- 
**Monitor relentlessly:** Visibility prevents surprises and enables optimization
- 
**Graceful degradation > hard failure:** Build systems that slow down rather than crash
- 
**Experiment to find your sweet spot:** Your workload patterns are unique

### 
  
  
  Recommended Starting Configuration:

- 
**Ollama:** Qwen2.5-Coder:14b-q4_K_M (~7.5GB) or Llama3.1:8b-q4_K_M (~4.5GB)
- 
**ComfyUI:** SD 1.5 with attention slicing and VAE tiling enabled
- 
**Scheduling:** Time-based with priority preemption for urgent requests
- 
**Monitoring:** Basic VRAM usage alerts + application health checks
- 
**Workflow:** Batch similar tasks together to minimize context switching costs

### 
  
  
  Final Thought:

The goal isn't to run both systems at maximum capacity simultaneously — that's unrealistic on consumer hardware. The goal is to have both systems available when you need them, with predictable performance and minimal frustration. With thoughtful scheduling, appropriate model choices, and basic monitoring, your RTX 5060 Ti can serve as a capable dual-purpose AI workstation for both development and creative work.

*Originally published on [ayraix.com](https://ayraix.com/signal/community/ollama-gpu-scheduling-rtx-comfyui-2026/), practical AI for enterprise builders.*
