Remix.run Logo
▲ manquer 2 hours ago

Context needs to pre-filled into a GPU memory in a node (usually 8xB300 or 8xH200) so there isn't any context or cache sharing between model families given their different parameter sizes, tokenizers, unlikely they are co-located in the same node.

Sub-agents not sharing context is a useful design-pattern when you want adversarial or independent reviews.

Cache reads could be shared between sub-agents, A single node(8GPU cluster) supports few hundred concurrent user sessions, that all share the same KV cache memory, so it is likely model providers do colocate your sub-agents in one node, it is more efficient , but may not be guaranteed so performance could vary; like we have with elastic compute and storage[1]

This can be cheaper depending on your coding flow i.e. cache hit % and the billing plan - cache reads are basically free or charged very little in subscription plans.

[1] Modern AWS does offer collocation at additional costs for compute but that is not the default and most other clouds do not offer it