The Agent Wars Go Local: Ollama, Unsloth, and Claude Code Battle for Your Desktop

Tags
digest
agents
local-ai
claude-code
ollama
AI summary
Published
September 29, 2026
Author
cuong.day Smart Digest
โšก
TLDR: The battle for your local AI agent stack is officially on. Ollama shipped a new System One API for structured decision-making, Unsloth dropped Laya Decision Models for local reasoning agents, and Claude Code upgraded to Sonnet 5.5 with a massive 1M context window. The message is clear: autonomous agents are moving from cloud demos to your machine.
Forget the cloud API wars for a second. Today's most interesting developments are happening on your desktop. We're seeing a coordinated push to make local AI agents not just possible, but *capable* of complex, structured reasoning. Ollama and Unsloth are giving developers the primitives to build agents that can make decisions, not just generate text. Meanwhile, Claude Code is betting that a 1M context window and a smarter default model (Sonnet 5.5) are what you need to manage sprawling codebases autonomously. This isn't incremental improvement - it's a fundamental shift in where and how we'll build the next generation of AI tools.

Is Your Local AI Agent Ready to Make Decisions?

The biggest story today isn't a single tool, but a convergence. Ollama v0.35.0 introduced the System One API, a new interface specifically designed for structured decision-making. It lets you give an agent choices, assign probabilities, and get back scored outcomes. This is huge because it moves local models beyond simple Q&A into the realm of actual reasoning and planning.
๐Ÿง 
What is the System One API? It's a new Ollama endpoint that treats decision-making as a first-class citizen. Instead of just asking "what should I do?", you can present a model with a set of options, each with context, and get back a structured, probabilistic recommendation. This is the foundation for building reliable, local-first agents that can navigate complex workflows.
Not to be outdone, Unsloth v0.1.900-beta shipped Laya Decision Models. These are purpose-built models designed for local deployment of reasoning agents. The release also boasts a 4.5x speedup for image and video generation on Apple Silicon, making the local agent experience faster and more responsive. The trend is unmistakable: the tooling is catching up to the ambition of creating truly autonomous, on-device AI.
  • Ollama v0.35.0 โ€” System One API for structured choices and probabilities.
  • Unsloth v0.1.900-beta โ€” Laya Decision Models for local reasoning agents.
  • Claude Code v2.1.284 โ€” Sonnet 5.5 default, 1M context window, but watch for sandbox regressions.
  • LiteLLM v1.104.0-rc.1 โ€” Signed Docker images and per-second pricing for secure, transparent orchestration.

The AI Coding Agent Stack: Stabilizing or Fragmenting?

While the local reasoning primitives are being built, the coding agent ecosystem is in a fascinating state of flux. Claude Code made a bold move by making Sonnet 5.5 its default model, banking on its improved reasoning and massive context. However, the release notes flag a critical regression in sandbox behavior causing session freezes - a stark reminder that power and stability are often at odds.
The competition is heating up. OpenAI Codex released a stable rust-v0.158.0, focusing on polish like copy-on-select and MCP server support. Gemini CLI is in rapid nightly iteration with a security-first focus. Meanwhile, GitHub Copilot CLI shows signs of stagnation with low PR activity despite high issue volume, and OpenCode is emphasizing network resilience and multi-provider support. The landscape is fragmenting, with each tool carving a niche: Claude Code for raw power, Codex for stability, Gemini for security.

๐Ÿ“Š Tool | Latest Move | Strategic Bet

  • **Claude Code** โ€” v2.1.284, Sonnet 5.5 default โ€” Massive context + reasoning power
  • **OpenAI Codex** โ€” rust-v0.158.0 stable โ€” Desktop stability & MCP integration
  • **Gemini CLI** โ€” v0.63.0-nightly โ€” Security-first, policy enforcement
  • **OpenCode** โ€” v1.18.33 stable โ€” Network resilience & multi-provider
  • **GitHub Copilot CLI** โ€” v1.0.90-1 fixes โ€” High issue volume, low dev momentum
โš ๏ธ
Watch Out: Claude Code's sandbox regression is a serious issue. If you're relying on it for automated workflows, test thoroughly before upgrading. The power of Sonnet 5.5 is tempting, but stability is non-negotiable for production use.

Memory, Retrieval, and the Infrastructure for Persistent Agents

An agent that can't remember is just a fancy autocomplete. Today's news shows a massive push to solve the memory and retrieval problem, which is the key to making agents truly useful over time. Cognee is an open-source AI memory platform enabling persistent long-term memory for agents, even with small models. claude-mem is a persistent context layer that compresses session history and injects relevant context.
On the retrieval front, PageIndex is a paradigm shift, offering a vectorless, reasoning-based RAG with 97% storage savings. LEANN, the MLsys2026 Best Paper winner, enables RAG on everything with minimal storage, ideal for personal devices. Meanwhile, ragflow continues to be a leading open-source RAG engine fusing retrieval with agent capabilities for enterprise apps. The stack for building agents that learn and remember is maturing rapidly.
  • Cognee โ€” Open-source memory platform for persistent agent context.
  • claude-mem โ€” Compresses session history for relevant context injection.
  • PageIndex โ€” Vectorless RAG with 97% storage savings and high accuracy.
  • LEANN โ€” MLsys2026 Best Paper for minimal-storage RAG on personal devices.
  • ragflow โ€” Enterprise-grade open-source RAG engine with agent capabilities.

โšก Quick Bites: The Rest of Today's AI News

  • Anthropic IPO Prospectus reveals massive R&D spending and ambitious long-term goals, fueling debate on sustainable AI business models.
  • Infosys partners with Anthropic to co-develop AI agents for regulated sectors like telecom and finance, expanding into India with the Topaz platform.
  • Hindsight, Paperclip, VoiceStudio lead trending lists for agent-centric tooling, signaling explosive momentum in autonomous AI workflows.
  • AnythingLLM emphasizes privacy, persistence, and self-hosting for AI agents, key to the local-first trend.
  • affaan-m/ECC is an agent harness cutting token usage by up to 65% across Claude Code and other models.
  • NousResearch/hermes-agent is an evolving agent that grows with user needs, representing next-gen personal AI companionship.
  • DeepSeek V4.1 gets full ROCm (gfx950) support in vLLM with optimizations for MXFP4 sparse indexing.
  • Qwen3-VL-Embedding support for multimodal embeddings via OpenAI-style content arrays in llama.cpp.
  • vLLM and SGLang continue advancing distributed inference for scalable agentic workloads.
  • llama.cpp stabilizes speculative decoding and multimodal support with fixes for GCC 15 and Vulkan.
  • World Labs joins AMD to strengthen AI infrastructure with deeper hardware-software alignment.
  • Nvidia Watchdog Chip proposed to monitor AI agent behavior in real time, sparking debate on surveillance and privacy.
  • Project Swap research paper shows autonomous AI agents negotiating in micro-markets, with model quality outperforming instruction design.
  • RRSI is a Google Research framework for safe, iterative agent self-improvement, seen as promising but under-discussed.
  • MicroLLM Lab is a browser-based playground for sub-1GB LLMs, democratizing access to small models.
  • ESP32S3 LLM Cluster runs a 1.58-bit model on an ESP32S3 cluster, a milestone in edge AI.
  • AI-Assisted Code Rewriting using AI to convert vulnerable C/C++ to memory-safe Rust at scale, a potential game-changer for security.
  • Apple is researching homomorphic encryption for privacy-preserving machine learning.
  • Google faces criticism over corporate culture, ethics, and autonomy in AI research.
  • AI labs called for public investigation and scrutiny for accountability and democratic oversight.

โ“ FAQ: Today's AI News Explained

  • Q: What is Ollama's System One API? โ€” It's a new endpoint in Ollama v0.35.0 designed for structured decision-making. You can present a model with choices and probabilities, and it returns scored outcomes, enabling more reliable local AI agents.
  • Q: Why is Claude Code's Sonnet 5.5 upgrade significant? โ€” It makes Sonnet 5.5 the default model with a 1M context window, offering dramatically improved reasoning for large codebases. However, a critical sandbox regression causing session freezes has been flagged.
  • Q: What are Laya Decision Models in Unsloth? โ€” They are native models in Unsloth v0.1.900-beta designed for local deployment of reasoning agents. The release also includes a 4.5x speedup for image/video generation on Apple Silicon.
  • Q: How is the AI coding agent stack changing? โ€” It's fragmenting. Claude Code bets on power (Sonnet 5.5), OpenAI Codex on stability, Gemini CLI on security, and OpenCode on multi-provider resilience. GitHub Copilot CLI shows signs of stagnation.
  • Q: What's the big trend in AI agent infrastructure? โ€” The shift to local-first, persistent agents. Tools like Cognee, claude-mem, PageIndex, and LEANN are solving the memory and retrieval problems to make agents useful over time on personal devices.
  • Q: What's the significance of the Anthropic-Infosys partnership? โ€” It's a major push to bring AI agents into regulated industries like telecom and finance, using Anthropic's models and Infosys's Topaz platform for governance and compliance.
๐Ÿ”ฎ Editor's Take: Today's news isn't about bigger models - it's about smarter *primitives*. The System One API and Laya Decision Models are the building blocks for agents that can actually reason and plan locally. The cloud giants are fighting over enterprise contracts, but the real revolution is happening on your laptop. The first developer to combine Ollama's decision API with Unsloth's speed and a persistent memory layer like Cognee will build something truly transformative. The local agent wars have just begun.