← yayster.com

Local AI Debugging · June 2026

Why Your Local LLM Keeps Cutting Off

Two bugs in Hermes Agent's Ollama integration were silently capping output at ~860 tokens. Here's the full story and the fix.

* * *

Eight hundred and sixty tokens.

That's all I got. I asked for a detailed thirty-five-hundred-word explanation of quantum computing — got three paragraphs and a cliff edge.

* * *

The Problem

It started with a Reddit post. Someone running Hermes Agent with Ollama was getting responses chopped at around 860 tokens. Consistently. They'd set num_predict to -1 in their Modelfile. Didn't help. They'd crank max_tokens in the config. Nothing.

I was hitting the same wall. This isn't a resource problem. Ryzen 9 99500. Four GPUs — dual RTX 3060 Ti, a GTX 1070, and a GTX 975. The model was laguna-xs.2 quantized at q4_K_M — a 128K context GGUF model built for long-form output.

And every time I asked it to write something substantial, it would just… stop.

Look at the numbers. 131,000 tokens of context available. The model used 32,000. But completions? One thousand one hundred nineteen. The finish_reason was always length. Not stop. Not an EOS token. The system was deliberately cutting it off.

The worst part? It would then enter a truncation loop — trying to continue, getting cut off again, eventually triggering context compression. It was eating my context window alive.

The settings you configured were never reaching the inference engine.

The Investigation

The architecture is straightforward: your prompt goes to Hermes Agent, which calls Ollama's API on port 11434, which runs inference on the GGUF model. Somewhere in that chain, a token limit was being imposed that I didn't set.

First suspect: Ollama itself. But Ollama respects num_predict: -1. If you curl the API directly with that flag, you get unlimited output. So Ollama wasn't the problem.

And here's where it gets interesting. I wasn't debugging alone. I had Hermes itself — the agent — helping investigate its own bug. And at one point, Grok jumped in too. Multiple AI agents collaborating to fix an AI agent framework. That's the world we're living in now.

The agent delegated the investigation to itself. It found that HERMES_MAX_TOKENS was being set, but the value wasn't propagating to the Ollama API call. We were getting closer.

Then I curled abzu — that's my Ollama server — hit the /api/show endpoint to see what the model was actually reporting.

{
  "details": {
    "families": ["laguna"]
    // MISSING: context_length at details level
  },
  "model_info": {
    "laguna.context_length": 32000,
    "laguna.rope.scaling.original_context_length": 32000
  }
}

And there it was. The context length — 32,000 tokens — was right there in the metadata. But it wasn't under context_length. It was under laguna.context_length. Namespaced by model family. Buried inside model_info, not details.

The parser in model_metadata.py had a hardcoded list of keys to check: context_length, context_window, max_model_len. None of them matched. So the system fell all the way through every detection path and landed on a fallback default of 2,048 tokens.

2,048 tokens. That's about 3,000 words. That's why everything was getting chopped.

But that was only half the story.

The Real Culprit

The primary bug was in chat_completions.py — the file that actually constructs the API call to Ollama.

When Hermes builds the request payload, it's supposed to pass through your max_tokens setting. But for Ollama-hosted models, it wasn't applying the override. Even if you set num_predict: -1 in your config, the code wasn't propagating it to the actual API call. It was using some internal default — and that default was low.

I applied the fix, restarted Hermes, ran the same test. Still truncated.

That moment where you think you fixed it and it's still broken? Yeah.

That's when I found the second bug. model_metadata.py needed wildcard pattern matching. The function _extract_first_int was doing exact key lookups. But GGUF models namespace everything — laguna.context_length, llama.context_length, qwen.context_length. I wrote a new function, _is_gguf_family_match, that checks for the {family}.{suffix} pattern across all model families.

Two files. Two focused fixes. That's it.

Fix 1: chat_completions.py

For any model endpoint on :11434 (Ollama's default port), force num_predict = -1 and set a max_tokens floor of 16,384. Eight lines of code.

# Ollama-specific handling for all models using local Ollama backend.
_base_url_str = str(params.get("base_url") or "")
if ":11434" in _base_url_str:  # Local Ollama default port
    api_kwargs["max_tokens"] = max(
        api_kwargs.get("max_tokens", 0), 16384
    )
    extra_body["num_predict"] = -1

Fix 2: model_metadata.py

Added wildcard pattern matching for {family}.context_length across all known model families: laguna, llama, qwen, gemma, and others. The new function automatically detects any namespaced context length key without hardcoding each family.

def _is_gguf_family_match(key: str, base_suffix: str) -> bool:
    """Check if key matches {family}.{suffix} pattern for GGUF models."""
    key_lower = str(key).lower()
    expected_suffix = f".{base_suffix.lower()}"
    return '.' in key_lower and key_lower.endswith(expected_suffix)

The Result

Same model. Same prompt. Same hardware. Twenty-one thousand context tokens. The response just kept going — quantum computing, superposition, entanglement, real-world applications — all the way to a natural conclusion. No truncation loops, no context compression, no fighting the system.

finish_reason is now stop — meaning the model decided it was done, not the infrastructure. That's how it should always work.

Full QA pass. Output length regression — passed. GGUF pattern matching — passed with one minor edge case. Context detection — passed. 122 pytest tests confirming it. And across fourteen sessions after the fix, the model generated 181,019 output tokens without a single truncation.

1,119 tokens → 21,167 tokens. Same model, same hardware. ~19x improvement.

The Mess Nobody Shows You

The session crashed on me twice. The terminal got nuked mid-investigation. I had to restart, re-delegate, re-verify. At one point the agent was helping debug its own framework while coordinating with another AI model on a separate machine.

That's the real story of local AI development. It's messy. But you push through because when it works, it works on your hardware, your terms, no API bills.

What This Means

One edge case remains — models that store context length under rope.scaling.original_context_length. Low impact, but worth noting if you're running something exotic.

The bigger lesson is about local AI infrastructure. When you run models locally, you own the entire stack. That's powerful — but it also means bugs hide at integration boundaries. The model was fine. Ollama was fine. The agent framework had a gap in how it talked to the backend. Nobody's at fault — it's just what happens when you stack open-source components.

This is why I document these things. Not because the fix is complicated — it's not. But because finding it requires understanding the full chain. And if you're building with local agents, you'll hit boundary bugs like this. Knowing where to look is half the battle.

Technical Details

Root Cause Analysis

IssueRoot CauseImpact
Output truncationnum_predict not passed to Ollama APICritical
Context detectionNamespaced GGUF keys not matchedMedium
rope.scaling edge caseWildcard doesn't match nested keysLow

GGUF Pattern Matching Test Results

Key PatternStatus
laguna.context_lengthDetected
llama.context_lengthDetected
qwen.context_lengthDetected
gemma.context_lengthDetected
details.context_length (legacy)Detected
laguna.rope.scaling.original_context_lengthEdge case

Hardware

Resources

Files Changed

Hit this bug on your own setup? Tell us what you're seeing — a human reads every message.

This is Yayster. Stoicism, Sumerian history, and the things you learn when you build with local AI agents. Every post documents something real.

Reach us on Telegram → t.me/yaysterllc_official

YouTube · X/Twitter