AetherGate context benchmark

Five fictional workloads, two product arms, three replicates. We publish outcomes including failures. No savings or accuracy claim is made until the complete matrix has been measured and graded.

Baseline uses one model; Foreman uses a two-model workflow. Costs are AetherGate catalog estimates from reported usage, not provider invoices. Synthetic operational review notes expand the original authored corpus. Provider automatic caching and unsupported temperature controls are limitations.