agent discoverability
Coding agents now decide which tools get installed. Armature helps your product get found and chosen by Claude Code and Codex.
why this is new
Traditional GEO measures how chat assistants cite your brand. Coding agents make the installation decision themselves, and the decision depends on the repository they work in.
Chat assistants · Claude, ChatGPT, Perplexity
Coding agents · Claude Code, Codex, Cursor
The same product wins in one repository and loses in the next. The repository shapes the queries, and the queries shape the choice. Armature is built for this surface.
1 · repository context
Language, libraries, docs and .md files
2 · agent priors
What the model already knows and trusts
3 · search queries
What the agent looks up on the web
4 · the pick
What gets installed in the code
what we do
We run any coding agent like Claude Code, Codex, Cursor or OpenCode inside a panel of repositories that match your category. You see when agents pick you, when they pick a competitor, and why.
We test every candidate change on real agent runs before it ships. The runs happen in sandboxes, never on your production docs. Only changes that improve the results go live.
We write the blog posts, docs, SDK guides, templates and listings. A growth engineer reviews every piece, and your team approves every change before it ships.
Every month you get your ranking movements, with the recorded agent runs behind each claim.
Each repository in the panel is written as a person who could buy your product, and mapped to a sector. We select the personas and sectors that match your buyers, so every run answers a question about your real market.
A dedicated growth engineer runs this process with you.
the first weeks
The first month builds your baseline and ships the first improvements. After that the loop repeats: experiment, ship, measure, report. The impact compounds from loop to loop.
We run the baseline panel on your category. You see your pick rate, your competitors and the reason behind each pick.
The first post and the highest-impact fixes ship, tested on agent runs and approved by your team.
We run the panel again and measure the movement against your baseline.
You get the first report with the recorded runs, and we queue the next experiments together.
New experiments and new content every week, and a report with evidence every month. Each loop builds on the last one, so the impact grows faster over time.
human in the loop
Every piece of content goes through the same pipeline, and you can inspect every stage.
Real agent runs show where you lose picks. We also monitor social channels and traffic data to catch the content trends that drive attention in your category.
Our agents draft the post, guide or listing with your product and your docs as context.
We score each draft with our own framework: coding agents run against a modified web that already contains the draft. We ship only the pieces that show positive signals.
An internal growth engineer reads, edits and signs off each piece before you see it. Nothing ships from an agent alone.
You see the final diff and approve it. Only then does the change reach your docs or your blog.
Every experiment runs in a sandbox, never on your production docs. Every claim in a report links to a recorded run. You can open any of them, at any time.
why not an agency
Agencies and GEO tools optimize the chat surface. They do not know which queries coding agents run inside real repositories.
The Armature way
The current default
leaderboards
We run coding agents inside our panel of repositories and record which tools they pick. We will publish the results here, one leaderboard per tool category.
Each leaderboard answers one concrete question, like "which error monitoring tool does Claude Code install in a legacy Django monolith".
Coming soon
Methodology: a panel of 22 representative repositories · several agents, models and prompts · every combination runs at least 3 times · custom experiments on request.
publications
We run continuous experiments on agent discoverability and publish the findings here.
Part I · How the selection works
Part II · What a vendor should do
Armature Search® mimics how Claude Code and Codex search the web. It reaches over 90% similarity with their results in our tests, so we can evaluate every change before it ships.
also from Armature
See the sessions users have with your product through Claude, ChatGPT and every other agent.
Run eval suites against your MCP (Model Context Protocol) server and catch regressions before you ship.
Tell us about your product. We will show you where you stand today and what we would change first.
or write to contact@armature.tech