跳至主要內容

SP Agent Team Token Report — Week of 2026-08-30

• 1 分
---
title: "Token Optimization Report: The Opus Spike and Hermes Offloading"
date: "2026-08-31"
category: "Engineering"
---

## This Week's Numbers

The latest telemetry for the period of August 24th to August 30th shows a significant regression in our token efficiency. We are currently operating under a **🔴 Red Light** status.

| Metric | This Week | Last Week | Trend |
| :--- | :--- | :--- | :--- |
| **Avg Opus Usage** | 64% | 24% | ↑ 40% (Worsening) |
| **Total Turns Added** | 1,991 | - | - |
| **Opus Turns** | 4,531 (cumulative) | - | - |
| **Sonnet Turns** | 2,435 (cumulative) | - | - |
| **Target Opus%** | < 50% | < 50% | ❌ Failed |

## What Changed

We observed a drastic shift in model dispatch. While the week started with a relatively balanced mix (61% Opus on 08-24), the latter half of the week saw an almost total reliance on Claude Opus. Specifically, between August 26th and 29th, Opus usage peaked at **97% to 100%**. This suggests a "complexity trap" where the agent defaulted to the most capable model for routine implementation turns rather than leveraging Sonnet's speed and cost-efficiency.

## Wins

Despite the Claude regression, the **Hermes Agent** (our secondary orchestration layer) is showing exceptional utility:
- **High Throughput:** Handled 65 sessions and 1,807 messages.
- **Heavy Lifting:** Processed ~125 million tokens, primarily using `deepseek-v4-flash-vision-exp`.
- **Tool Mastery:** Successfully executed 798 tool calls, with the `terminal` (43.1%) and `read_file` (18.8%) being the primary drivers.

## Challenges

The primary challenge is **Model Over-provisioning**. We are using Opus for turns that do not require its reasoning depth. The data shows that on August 28th alone, we clocked 1,540 Opus turns compared to just 8 Sonnet turns. This inefficiency creates a bottleneck in both cost and latency.

## Next Week's Target

**Objective: Return to Green Light Status.**
- **Primary Goal:** Reduce weekly average Opus% to **< 50%**.
- **Secondary Goal:** Increase the ratio of Sonnet turns during the implementation phase of the development cycle.

## Dispatch Optimization

Based on the Hermes utilization data, there is a clear opportunity to shift "Environmental Exploration" tasks away from Claude. 

Next week, we will shift the following to Hermes:
1. **File System Scanning:** Move all `search_files` and `read_file` loops to Hermes.
2. **Initial Shell Probing:** Use Hermes for `terminal` commands to verify environment states before invoking Claude for high-level logic.
3. **Basic Refactoring:** Offload routine boilerplate updates to `deepseek-v4-pro` via Hermes.

## Cost Savings

The cost delta between using Hermes (powered by DeepSeek) and Claude Sonnet is staggering. 

- **Hermes Total Spend:** ~$0.35 for 125M tokens.
- **Estimated Sonnet Cost:** At an average blended rate of ~$5.00 per million tokens, the same volume would have cost approximately **$625.00**.
- **Weekly Savings:** By routing these 1.8k messages through Hermes, we saved roughly **$624.65**.

## Recommendations

Based on this week's telemetry, we are implementing the following action items:

1. **Route more implementation to Sonnet**: Since Opus% (64%) is well above the 50% threshold, we will enforce a "Sonnet-first" policy for all coding tasks.
2. **Use free engines more**: Gemini usage was negligible (only 1 session). We will integrate Gemini-3.6-Flash into the research pipeline to reduce the burden on paid models.
3. **Review Dispatch Hooks**: Analyze the logic that triggered the 100% Opus streak on Aug 26-29 to prevent similar regressions.