Show HN: Open-source model routing for coding agents at Astra-level performance

A few months ago we started building a model router for coding agents because we thought we could outperform any single model with an ensemble approach. Recently we’ve achieved that milestone and I want to talk about how we did it.

First of all, a quick explanation: the Weave Router (https://github.com/weave-os/router) plugs into any coding agent (e.g. Claude Code or Codex) and intelligently switches between LLMs. So, for example, Astra handles tricky debugging or complex system design tasks, and Deepseek v4 Flash handles simple frontend updates.

What we’re announcing today is our new routing model, which we’re calling Weave Router 2.0. We benchmarked 2.0 against GPT-6 Astra on Terminal Bench 4.0 and SWE Atlas. On both benchmarks, the router had equivalent pass rates. On Terminal Bench, the router hit 52% of Astra’s cost, and completed tasks 2.2x faster. On SWE Atlas, the router cost 54% as much as Astra and ran 2.5x faster. (Full results on our website at https://weaveos.com/router!)

It turns out training a model to route effectively - taking into consideration model capabilities, costs, cache awareness, and more - is a really hard problem! I want to talk about three ways we were able to improve so much over the last few months: 1) a new architecture, 2) larger training data set size, and 3) smarter cache-eviction impact calculation.

1) a new architecture. Our initial approach used an RL model without many priors. While RL is still an important part of the story, the cost of fully exploring the space of routing decisions is very high, so we’ve taken some shortcuts that have significantly improved performance.

Consider how large the search space for the routing problem is. Take a typical coding agent session, with ~100 agent turns (i.e. 100 LLM API calls). Technically there are 100 chances to select a model. If we assume a roster of ~10 models (of course there are lots more but we can remove any that are Pareto dominated), then there are 10^100 possible paths through that session. We simply cannot explore all of them! So that's why clever tricks to shrink this space are so important.

In particular: we trained a hidden Markov model to trace the session state, then a classifier maps the session to one of a few buckets of similar models. Using the HMM allows us to evaluate not just where a session is currently, but how it got there. We've gotten significantly better performance on bucket selection by incorporating that information - we believe this is because two sessions that might look quite similar to a naive classifier are much better distinguished by this HMM approach.

Using this HMM + classifier to select a bucket first significantly shrinks the space to explore, by throwing out most models that could not reasonably serve the given session. This rearchitecture was the single biggest performance unlock!

2) larger training data set size (much less technically interesting but still an important part of the story). By using frontier LLMs to help us label a larger and more diverse set of coding agent sessions, we were able to bootstrap the two models discussed in 1) to a better state, while also providing even richer reward signals for RL.

3) smarter cache-eviction impact calculation. One of the hardest parts of routing well (if you care about saving money) is using the model caches intelligently. We built a subsystem that can calculate the expected value of switching models (and thus paying a high one-time cost to fill up a different cache) much more accurately, helping us avoid costly and unnecessary switches in more cases, while still switching when the benefit outweighs the cost. This is where most of our improvement on cost has come from.

We still have a lot of room to continue to improve (we won’t rest until we’re consistently beating Astra/Fable, not just tying!) but matching frontier model performance was a huge milestone for our routing model, and in my opinion validates our initial hypothesis that an ensemble of models can do better than any single model ever could.

Our router is open source (https://github.com/weave-os/router) so anyone can try it out. Or if you prefer you can use our hosted version (https://weaveos.com/router).

66 points | by adchurch 1 day ago

14 comments

  • YuechenLi 1 hour ago
    Have you compared this to using GPT-6.1 Sol instead of GPT 6 Astra + Deepseek? From my test, 6.1 Sol is a lot more token efficient than 6 Sol while being similar to Astra in performance, and I don't really find 6 Astra to be significantly better than 6/6.1 Sol for general coding as I feel 6 Astra is only noticeably better at spatial reasoning/vision compared to 6 Sol, and 6.1 Sol really closed the gap on that front.
  • svnt 38 minutes ago
    If you are training on data labeled by frontier models, how do you expect to exceed the performance of frontier models, other than in the cost dimension by recognizing simpler problems and routing to cheaper models?
    • adchurch 15 minutes ago
      Different frontier models are good at different things! We'll be the ones combining them optimally.
  • gitowiec 38 minutes ago
    Can it route to locally or LAN hosted Qwen or some other open weights model?
    • adchurch 10 minutes ago
      Not yet! Something we're interested in experimenting with though, hit me up at andrew@weaveos.com if you have any thoughts here
  • ajspig1 1 hour ago
    How do you handle provider variance on OpenRouter for the opensource models? Or do you use your own hosted version to mitigate this?

    And for both opensource and closed source, does the router account for provider quality, or catch it when a provider degrades?

    • adchurch 11 minutes ago
      For our hosted version we carefully select which providers we use (we don't use OpenRouter). For the self-hosted version though, OpenRouter does make it much easier to get started (at the cost of that variance potentially affecting quality).

      Yes for both! We have some logic to put providers on cooldowns and/or deprioritize them.

  • rirze 2 hours ago
    How does this choose which models to use with any arbitrary set of model providers to work from? And why is an openrouter necessary for self-hosting?
    • adchurch 14 minutes ago
      To the point of using buckets of models: as long as there's >0 models available in a bucket, and we can order models in order of fit, we're resilient to different sets of models being available. With that said obviously cutting out some models has a much larger effect than others.

      OpenRouter isn't strictly necessary but it does make it easier to not have to set up accounts/API keys with several different providers to get started.

  • jamesforestwest 2 hours ago
    How do you define the model buckets, and what happens when a session genuinely needs a model that isn't in the bucket the HMM picked?
    • adchurch 23 minutes ago
      You can think of buckets as models with similar capabilities. So for example Deepseek 4.1 Flash will not be in the same bucket as Astra.

      The latter is an interesting question! In practice because the session continues, we can see it's going down the wrong path and escalate. Basically no decision we make when routing is entirely unsalvageable (but we do have a performance penalty for every incorrect decision we make so of course we try to avoid it).

  • thefourthchime 3 hours ago
    Interesting work, and thanks for describing how your router works internally. It's definitely a fascinating subject. How would you say this compares to Cursor's auto mode?
    • adchurch 2 hours ago
      Absolutely!

      Conceptually very similar to Cursor's auto mode. The key distinctions are:

      - We plug into any harness (e.g. Claude Code, Codex, OpenCode, Pi)

      - We aren't incentivized to route to our own model, we're incentivized to route to the best model whatever it may be

    • aschla 3 hours ago
      And similarly, Copilot’s Auto mode?
      • adchurch 16 minutes ago
        I haven't tried it as recently but last time I checked they only route once per session (or subagent). Imo this is basically impossible to do correctly. Consider the case where you start with one prompt "rewrite this in rust". Trivial in a 1 day old repo, extremely difficult in e.g. the VSCode repo!
  • 1minusp 2 hours ago
    Does this allow for a predefined budget?
  • redrove 4 hours ago
    Is the model you trained available as open weights?
  • aminsamir45 2 hours ago
    AGI is here!
  • mrkn1 33 minutes ago
    [flagged]
  • itsmeduncan 3 hours ago
    [flagged]
  • xms17189 20 hours ago
    [dead]
  • adurnos 2 hours ago
    [dead]