Remix.run Logo
▲ heyjstn 2 hours ago

Have anyone tried a workflow that:

- Fable 5.1 for planning/adversarial reviewer

- Opus 5.5 for well-scoped tasks break down

- Sonnet 5.5 for these well-scoped tasks implementation

I think the blocker might be how efficient the context is compacted and sending around between these agents

▲afro88 2 hours ago | parent | next [-]

Opus 5.5 in my experience outshines Fable 5.1 anyway. May as well have Opus do plan, breakdown and review, and Sonnet implement.

▲chrismustcode 2 hours ago | parent | prev | next [-]

You might as well use Opus for everything there.

Changing model would be cache busting spiking usage for no good reason when Opus can do it all.

Haiku 5.5 might fit well though depending on pricing.

▲SirMadam 2 hours ago | parent [-]

Do subagents share context? If Opus delegates to a different Sonnet window, I don't believe this busts cache?

▲manquer an hour ago | parent | next [-]

Context needs to pre-filled into a GPU memory in a node (usually 8xB300 or 8xH200) so there isn't any context or cache sharing between model families given their different parameter sizes, tokenizers, unlikely they are co-located in the same node.

Sub-agents not sharing context is a useful design-pattern when you want adversarial or independent reviews.

Cache reads could be shared between sub-agents, A single node(8GPU cluster) supports few hundred concurrent user sessions, that all share the same KV cache memory, so it is likely model providers do colocate your sub-agents in one node, it is more efficient , but may not be guaranteed so performance could vary; like we have with elastic compute and storage[1]

This can be cheaper depending on your coding flow i.e. cache hit % and the billing plan - cache reads are basically free or charged very little in subscription plans.

[1] Modern AWS does offer collocation at additional costs for compute but that is not the default and most other clouds do not offer it

▲enraged_camel an hour ago | parent | prev [-]

Subagents don't share context. But that's why delegating implementation to a subagent doesn't work well except for things that are truly mechanical in nature: the subagent needs to independently reason about the task it is given, and then the output will also be reasoned about by the main agent. So you end up wasting time and tokens.

▲mnicky an hour ago | parent [-]

On the contrary, subagents save context overall, when the task is sufficiently large.

Also, my experience is that Fable 5.1 is very good at prompting/orchestrating Opus/Sonnet subagents when working on a larger task (e.g. 1-2M context window use only for the orchestrator itself).

▲Aboutplants an hour ago | parent | prev [-]

Do you even need Fable for much of anything now? I’m basically using it as a reviewer at the end of whatever I’m working on, and even then I’m really not finding much benefit.