TensorFlow Out of Memory Error: Causes, Fixes & GPU Memory Optimization
Look, you’ve probably stared at a black terminal, watched the GPU fan scream, and then got hit with that dreaded ResourceExhaustedError. I’ve been there at 2 am, coffee cold, wondering why my tiny model suddenly ate the whole GPU. The truth? OOM errors are sneaky, they love hiding in plain sight. Let’s rip that mystery open.
Understanding TensorFlow OOM Errors
TensorFlow talks to the GPU driver, asks for a chunk of memory, and boom nothing left. The error usually looks like this:
tensorflow.python.framework.errors_impl.ResourceExhaustedError: OOM when allocating tensor with shape[256,1024,1024,3] and type float32
That line is the system’s way of saying, “I ran out of RAM on the card, pal.” It isn’t just about the model size; it’s about the whole graph you built, the batch you feed, and the way TensorFlow pre‑allocates buffers.
Pro tip: TensorFlow by default grabs almost the entire GPU memory up front. That’s why you see OOM even before the first batch lands.
Common Causes of Out of Memory Errors
Batch size too big – the classic rookie mistake.
Unnecessary variables – you left a
tf.Variableinside a loop, creating a new tensor each step.Large input pipelines –
tf.datashuffling huge datasets in memory.Mixed‑precision mis‑config – you thought you were using FP16 but forgot to set the policy.
Multiple models in one process – serving and training together.
Think about it: you’re trying to squeeze a 10‑GB elephant into a 4‑GB suitcase. It won’t work.
Real‑world scenario
At my last gig, we were training a ResNet‑152 on 8 GB of VRAM to detect anomalies in X‑ray images. The model alone was 2 GB, but the batch of 32 1024×1024 images added another 4 GB. The OOM hit us right after the first epoch. The fix? Shrink the batch, enable gradient checkpointing, and let TensorFlow grow memory lazily. That saved us days of wasted GPU time.
Diagnosing Memory Usage
First, peek at what TensorFlow thinks is happening.
# Bash: show GPU memory per process
nvidia-smi
You’ll see something like:
+-----------------------------------------------------------------------------+
| GPU PID Process name GPU Memory |
| 0 12345 python 7983MiB / 8192MiB |
+-----------------------------------------------------------------------------+
If your script is hogging 8 GB before any training, you know the culprit is pre‑allocation.
Second, enable TensorFlow’s memory profiler.
import tensorflow as tf
tf.config.experimental.set_memory_growth(tf.config.list_physical_devices('GPU')[0], True)
Run the script, then look at the logs:
2026-10-02 10:15:33.123456: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1234] Memory growth enabled for GPU:0
If you still see OOM, dump a detailed memory snapshot:
tf.profiler.experimental.start('logdir')
# ... run training ...
tf.profiler.experimental.stop()
Open the Chrome trace viewer, filter by “GPU Memory”.
Warning: Don’t forget to stop the profiler, otherwise you’ll fill your disk.
Fixing OOM Errors: Code‑Level Strategies
What NOT to do
# Wrong: creates a new variable each loop iteration
for i in range(1000):
w = tf.Variable(tf.random.normal([1024, 1024]))
# ... use w ...
That code silently leaks memory. The error you’ll see:
ResourceExhaustedError: OOM when allocating tensor...
The fix
# Right: create once, reuse
w = tf.Variable(tf.random.normal([1024, 1024]))
for i in range(1000):
# ... use w ...
Now the graph reuses the same buffer; no leak.
Another common pitfall: oversized batch
# Wrong: batch = 128 on a 4GB card
train_dataset = ds.batch(128)
Error: OOM right at model.fit.
Fix:
# Right: start small, let TF auto‑scale
train_dataset = ds.batch(16)
If you need speed, enable tf.data.experimental.prefetch:
train_dataset = ds.batch(16).prefetch(tf.data.AUTOTUNE)
Gradient checkpointing (a.k.a. recompute)
When you have a deep net, you can trade compute for memory:
# Wrong: full graph kept in memory
model = tf.keras.applications.ResNet152(weights=None, input_shape=(224,224,3))
# Right: use tf.recompute_grad
@tf.function
def forward(x):
return tf.recompute_grad(model)(x)
# training loop stays the same
Yes, it slows down a bit, but you avoid the OOM without shrinking batch.
Optimizing GPU Memory Allocation
TensorFlow gives you two knobs: memory growth and virtual device limits.
Memory growth
gpus = tf.config.list_physical_devices('GPU')
if gpus:
try:
# Enable growth for each GPU
for gpu in gpus:
tf.config.experimental.set_memory_growth(gpu, True)
except RuntimeError as e:
print(e)
Now TensorFlow only grabs what it needs, leaving room for other processes.
Virtual devices
# Limit each GPU to 4GB
tf.config.experimental.set_virtual_device_configuration(
tf.config.list_physical_devices('GPU')[0],
[tf.config.experimental.VirtualDeviceConfiguration(memory_limit=4096)]
)
You can split a single GPU into “virtual” pieces, useful when you run inference alongside training.
Mixed precision
If your GPU supports FP16, enable it:
from tensorflow.keras import mixed_precision
mixed_precision.set_global_policy('mixed_float16')
Now weights stay FP32, activations go FP16, halving memory for most layers. Watch out for loss scaling, but Keras handles it automatically.
Clear session after heavy ops
# After a huge model is done
tf.keras.backend.clear_session()
That releases all TF graphs and frees the GPU memory back to the OS.
Good Habits for Efficient Model Training
Profile early – run a short epoch with
tf.profilerbefore scaling up.Use
tf.function– graph mode reduces overhead and memory fragmentation.Avoid
tf.concaton the GPU in a loop; accumulate on CPU then transfer.Cache static data –
dataset.cache()keeps it in RAM, not GPU.Prefer
tf.keras.layers.Lambdasparingly – they hide extra tensors.
Side note: I once wrapped a custom loss in a
Lambdalayer, thinking it was neat. It turned the loss into a separate tensor each step, blowing memory. Removing the Lambda fixed it instantly.
Conclusion
Out of memory isn’t a mystery; it’s a symptom of a graph that asks for more than the card can give. By shrinking batches, reusing variables, enabling memory growth, and sprinkling mixed precision, you can train big models on modest GPUs. The key is to look at what TensorFlow actually allocates, not what you assume.
I’ve seen the OOM monster bite my code at 2 am, then watched it retreat after a few lines of fix. Your next training run will be smoother if you follow the tricks above.
Quick Summary
Don’t pre‑allocate the whole GPU – turn on
set_memory_growth.Batch size matters – start tiny, scale up gradually.
Reuse variables – avoid leaks inside loops.
Mixed precision = memory win – set the global policy.
Profile early, clear sessions – keep the graph tidy.
I’ve used these tricks on a low‑end laptop for a BERT fine‑tuning job (see my post on TensorFlow GPU Memory Allocation Error Low‑End Laptop Fix). It turned a nightly crash into a smooth 3‑hour run.
FAQs
Q: Why does nvidia-smi show 100 % memory usage even before training starts?
A: TensorFlow pre‑allocates most of the GPU memory by default. Enable memory growth (set_memory_growth) to let it allocate on demand.
Q: Can I run two TensorFlow models on the same GPU simultaneously?
A: Yes, but you must limit each model with virtual device configs or use set_memory_growth. Otherwise they’ll fight for the same memory and cause OOM.
Q: How do I know if my dataset pipeline is eating GPU memory?
A: Check the TensorFlow profiler trace for “GPU Memory” spikes. If you see large buffers tied to tf.data ops, add .cache() on CPU or reduce prefetch size.
Q: Does tf.keras.Model.fit automatically free memory after each epoch?
A: Not completely. It reuses the same graph, so any tensors you create inside the training loop that aren’t tracked will linger. Use tf.keras.backend.clear_session() when you’re done with a model.
Q: Is mixed precision safe for all models?
A: Mostly, but some ops (like certain loss functions) can underflow in FP16. Stick to the mixed_float16 policy and let Keras handle loss scaling; test your validation loss to ensure it’s stable.



