Fixing SSE Buffer-Bloat in Local Tunnels: Guaranteeing Zero-Latency Streaming for Local LLMs
You’ve just deployed a state-of-the-art local LLM using vLLM , Ollama , or llama.cpp . When you test it on localhost , the token generation is a thing of beauty—a smooth, continuous stream of text that feels instantly responsive. But the moment you expose this endpoint to the outside world via a reverse proxy (like Nginx) or a local tunnel (like Cloudflare Tunnels or Ngrok), the magic dies. Instead of a smooth flow, your frontend receives nothing for several seconds, followed by a massive, chunky block of text all at once. If you are building real-time conversational AI interfaces, this "choppy" token output destroys the User Experience (UX). It ruins your Time to First Token (TTFT) metrics and makes your application feel sluggish, regardless of how fast your GPUs are actually inferencing. The culprit? SSE Buffer-Bloat . In this guide, we will dissect why standard networking layers inadvertently sabotage Server-Sent Events (SSE) and HTTP chunked transfer encoding. More ...