Skip to content

Agent tooling audit ​

An MCP client loads every tool definition from tools/list into the model's context before the first user message. This page measures that cost for levelrail-mcp, lists what is heavy or redundant, and describes the test that stops it growing unnoticed.

Numbers are estimates: the serialized tools/list JSON length divided by 4 (no tokenizer library is in go.mod). They are stable enough to compare modes and catch regressions, not to bill against.

This page complements MCP tool surface, which reports the per-toolset cost and the agent-core profile. The difference is scope: that report counts each tool's name, description and input schema (about 18,300 tokens for 144 tools), while this audit measures the whole serialized tools/list response a client actually receives, which also carries output schemas and annotations (about 59,800 tokens for the same tools). Output schemas are the part the surface report does not see, and they are 57% of the bytes. This page adds the per-mode numbers, the per-tool findings and the regression test.

Tools and tokens per mode ​

Measured by go test -run TestToolListTokenBudget -v ./internal/mcptools (or scripts/mcp-token-budget.sh).

ModeToolsEstimated tokens
read-only10941,600
standard (default)13555,600
full14459,800

Every mode is defined by internal/mcptools/modes.go and driven by the class table in internal/mcptools/classes.go: 109 read, 26 mutating, 9 destructive tools.

Where the tokens go (full mode) ​

Part of a tool definitionShare of bytes
Output schema57%
Input schema16%
Description15%
Annotations and _meta7%
Name and framing5%

The output schema is the dominant cost. Typed handlers derive it from the Go result struct, so AppResource (about 3,000 characters) is repeated in get_app, list_apps and restart_app, and the deploy result schema (about 3,900 characters) is repeated in deploy_app, rollback_app and approve_deploy_approval. A model rarely needs the schema to read a result that is also returned as JSON text.

By toolset ​

ToolsetToolsReadMutateDestructiveEst. tokens
deploys20135213,500
apps128316,200
models138325,500
alerts107304,100
loadbalancer95313,700
logs117403,300
backups118303,100
pipelines33002,400
domains109102,300
metrics44001,900
iac21011,800
nodes44001,800
diagnostics33001,400
everything else (13 toolsets)322930about 8,000

Heaviest tools ​

ToolEst. tokensOf which output schema
promote_app1,390963
deploy_app1,266963
rollback_app1,253963
approve_deploy_approval1,127983
clone_app1,106764
explain_pipeline_run1,076894
apply_resources1,075612
preview_promote_app1,003651
deploy_compose959806
preflight_model929614

Findings ​

Overlapping tools ​

ToolsOverlapRecommendation
get_app_status, list_deploysIdentical handler and result (current reconcile conditions).Drop list_deploys from agent-facing modes; keep for compatibility in full.
deploy_app, rollback_appSame request and handler.Keep both (intent is useful and the classes differ) but share one description.
list_deploys, list_deploy_attempts, list_deployments, list_failed_deploysFour ways to list deploy history at different scopes.Descriptions must say which scope (one app vs fleet, failures only). Prefer list_deploys or list_deploy_attempts in compact profiles, not both.
get_resource_recommendation, get_database_resource_recommendationSame shape for apps and databases.Fine, but one description should point at the other.
get_app_logs, get_model_logs, list_archived_logsThree log entry points.A compact, capped log query tool for agents is planned.
preview_promote_app and promote_app, preview_clone_environment and clone_environmentPreview and apply pairs.Correct as designed; the read-only preview is what an agent should call first.

Descriptions ​

  • 19 descriptions exceed 350 characters; the longest are diagnose_app_failure (726), promote_app (667), clone_environment (583), deploy_app (578), clone_app (575) and rollback_app (519).
  • 9 descriptions cite implementation details a model cannot use (cmd/levelrail-cli, internal/..., "the web dashboard's button"): diagnose_app_failure, get_attention, get_system_doctor, list_archived_logs, list_backup_targets, list_model_cache, prune_system, rollback_app, test_storage_destination.
  • Four descriptions are under 80 characters and do not say when to use the tool: get_database, restart_app, list_databases, clear_app_load_balancer.
  • Every input parameter already has a description (0 missing), so schema documentation is complete; the cost is repetition, not omission.

Annotations ​

  • Every registered tool has a title, readOnlyHint, openWorldHint and (for non-read tools) destructiveHint; TestEveryToolClassified fails the build otherwise, so none are missing.
  • idempotentHint is inferred from a set_ name prefix only. set_* tools are idempotent; approve_*, reject_*, expire_* and restart_app are also safe to repeat but are marked non-idempotent.
  • The title is a mechanical rewrite of the name ("Get app status") and adds tokens without adding meaning; clients already show the name.
  • Sensitive and untrusted-output tools carry _meta flags; that is working as intended.

Schema bloat ​

  • set_app_load_balancer (1,555 characters of input schema), plan_apply and apply_resources (1,254 each), deploy_model (1,166) and list_deployments (1,047) carry the largest input schemas. All are justified by real option surface, but they are poor candidates for a small agent-facing mode.
  • Nullable arrays and maps are rendered as "type": ["null","array"], which is longer than needed.

The agent-core profile ​

The profile lists 15 tools and omits output schemas from tools/list (results are still returned as text and structured content). Estimated cost of the whole listing:

ListingEst. tokens
Profile with output schemas (before)about 6,500
Profile without output schemas (now)about 2,500
full modeabout 60,500

get_app_env (plain env values plus secret key names, never values) and query_logs (capped, filtered log excerpt) were added to the profile; get_app_logs was replaced by query_logs. The profile's total listing is held under a budget by TestAgentCoreProfileListingBudget (default 3,500, override with APP_MCP_TOKEN_BUDGET_AGENT_CORE).

Recommendations ​

  1. Use the agent-core profile in MCP tool surface for autonomous agents (see the numbers below).
  2. Trim outputs before trimming inputs: return compact results from agent-facing tools rather than the full resource structs, which removes the largest output schemas from the list.
  3. Rewrite the 19 long descriptions to one sentence that says what the tool returns and when to call it; move history and CLI comparisons into docs.
  4. Mark restart_app, approve_*, reject_* and expire_* idempotent, and drop the mechanical title if a client ever charges for it.
  5. Keep the budget test below passing; raise a budget only with a reason in the PR.

Regression test ​

internal/mcptools/budget_test.go serializes tools/list for each mode over an in-memory MCP session and fails when the estimate exceeds the mode's budget.

Env varEffect
APP_MCP_TOKEN_BUDGET_READ_ONLYOverride the read-only ceiling (estimated tokens).
APP_MCP_TOKEN_BUDGET_STANDARDOverride the standard ceiling.
APP_MCP_TOKEN_BUDGET_FULLOverride the full ceiling.

Run it with the report of the heaviest tools:

bash
scripts/mcp-token-budget.sh

Released under the Apache 2.0 License.