Product

Collaborate: Improve AI agents like code, wherever you work

Collaborate lets operators, engineers, and AI agents work as peers, with version control, evals, and continual learning built in.

Photo of Emma Martin

Emma Martin

Photo of Neal Lathia

Neal Lathia

·

Blog cover image

Introducing: Collaborate

Today we're launching Collaborate: a new set of products that enables the large, multi-disciplinary, and regulated teams who manage Gradient Labs AI agents to improve them using the tools they already use every day. Our vision is to empower teams to quickly and continuously iterate on their AI agents, with any AI agents.

Collaborate has two parts. First, we’re releasing a new foundation for all of our AI agents that is fully version controlled, like code. Engineers can edit any part of our AI agents by directly drafting a pull request in a repository. Operators can make the same changes using our web product. No matter where you work, every change can be drafted, tested, reviewed, and approved before being launched, leaving behind a full audit trail.

Our new foundation also comes with a suite of tools and features that, collectively, enable you to continuously improve your AI agent, ensuring that changes are rigorously tested, regressions are caught before launch, and every misstep the agent takes becomes a case to actively test against in the future.

We have seen how much rigour financial services teams must apply when managing their AI agents. Engineering, operations, QA, and compliance do not need to work over one another, across documents and siloed tools. Collaborate removes that constraint, so your automation rate and customer satisfaction can keep climbing.

Here is what Collaborate looks like for each of the teams that shape your agent:

Collaborate, for Engineers

Scaling an AI agent past FAQ responses has always required engineering: wiring up the tools it calls, connecting the APIs behind them, and giving the AI agent the critical customer context that it needs.

Until now, tools and resources were primarily added to our platform via API; procedures and configuration were mostly updated via our web product. With Collaborate, you can build, test, and launch every change straight from a git repository.

A view of GitHub, for engineers to edit their agent directly as code

Your agent's procedures, tools, and test cases are now available as markdown files in a repository: they are all files that you can grep, diff, and change in bulk.

Any change to the agent's behaviour is a pull request: drafted and tested on a branch, reviewed by your team, and merged only when the checks pass. Multiple changes on the same procedure can be created and tested in parallel. Like code, merge conflicts must be resolved before merging, so edits can no longer overwrite each other. Critically, all changes now come with a version history, capturing who made what change (and why), which elevates your ability to audit any behaviour retrospectively.

Collaborate, for Product Managers and Operators

Collaborate brings all the rigour and auditability of code, without requiring less technical operators to know all the minutiae of version control. Operations and content teams get the same safeguards engineers rely on, straight from our web product, where they appear as branches, drafts, approvals, and instant undo.

An image showing review of a failed test

An operator can work on several changes at once without touching the live agent, and test each new version offline. When a change is ready, sign-off happens on the change itself, so one editing session leaves one clean change behind rather than twenty versions of the same procedure.

Before anything goes live, every test case runs against the change in one batch, and if something still makes it through, the history shows exactly what changed and who changed it, ready to roll back in one click. The result is a team that tackles more use cases, collaborates more easily with engineering, and satisfies compliance, all without writing a single line of code.

Collaborate, for AI agents

Many teams who have partnered with us are already drafting procedures with Claude, Codex, or their in-house AI tooling. Until now, that meant pasting the output back into the platform by hand, one procedure or test case at a time.

Now, any AI agent can directly improve the Gradient Labs agents. Ask Claude to draft a change to your dispute procedure and write three test cases for it, and it reads the current procedure, drafts the change on its own branch, and writes the test cases to prove it. The change shows up the way a colleague's would: on a branch, waiting for approval before merging.

You can also point any coding AI agent at your repository to turn any document into a whole suite of test cases. And to maintain safety and quality, every change an AI agent drafts is subject to the same approvals as your human team.

An image showing a Claude Code chat window, to edit the agent directly with Claude

Collaborate and self-improvement

We are moving away from treating agent procedures as documents, and re-imagining the entire experience as code. Just like engineers use code to deploy and manage the most complex systems in the world today, Gradient Labs agents can now be treated as code that continuously gets extended and improved. Here's how it works:

  • Build: Draft a change from the web app, your own repository, or ask the AI agents you already use to do it for you. Every path produces the same reviewable change.

  • Test: Validate the change against your test cases, and chat directly with the new version offline.

  • Review: Teammates review and comment on the change, the suite has to be green before it merges, and merging puts it live.

  • Launch: Test suites run at scale before launch, including pre-built safety checks like toxicity benchmarks. New capabilities can be rolled out gradually, taking a small share of live traffic before the full volume.

  • Monitor: Scorers watch live traffic, tracking the signals you care about, and firing alerts when quality dips.

  • Improve: The agent proposes fixes from what it sees in production, makes every failure a permanent test case to prevent future regressions, and nothing goes live without your approval. Improvements feed straight back into the Build phase.

A diagram which shows the 6 stages of self-improvement

The improve step is kickstarted autonomously, while still under your supervision: it analyses the conversations it escalated and proposes procedure changes, and it spots the knowledge it was missing by watching how your human agents answered.

Over time, your team's job becomes reviewing changes rather than writing them, and your customer operation moves closer to running on autopilot.

Test cases for QA and compliance

The testing suite in the Gradient Labs platform

For QA and compliance, every change to an agent raises the same questions: what changed, who approved it, and is it compliant? Collaborate keeps the answers to the first two on record by default, and puts the third in front of you before launch, with the change and its test results in one reviewable place.

Our test suites run scenario-based synthesised conversations. They can be based on a real conversation that has happened in the past, and simulate how the agent would continue from a given starting point, with pre-defined tool responses and resources. A green suite becomes the bar for launch, and you can trigger runs from your own CI.

That record turns an audit from an investigation into a lookup. When a regulator asks why the agent said what it said, the answer is on file: the version that was live, the change that introduced it, and the approval behind it.

A better experience for every customer

This cycle of self-improvement shows up in two places: a higher automation rate for your operations, and a better experience for your customers. To imagine what that looks like, take a customer who flags a charge they don't recognise. The agent verifies who they are, opens the case, and gets to work: it chases the missing evidence, submits the chargeback, and days later closes the case with the outcome, without the customer ever being handed off.

Now suppose the agent handled one of those steps incorrectly. That conversation becomes a test case, and the fix is proven against the exact conversation that failed before going live, so every customer after that gets the better version. Week by week, the agent covers more of the case and handles it more accurately, and no customer ever hits the same failure twice. Taken far enough, this points at something banks have never been able to offer at scale: every customer getting the experience of their own dedicated account manager.

The future of work

We believe in a future where anyone in financial services has the tools to one-shot production-grade agents. Getting there safely, as agents multiply across departments and use cases, will require operators, engineers, and AI agents to work as peers. With Collaborate, they already can.

If you believe in this future, or are struggling to maintain your own AI agent, book some time with us. We’d love to help.

Share post

Copy post link

Link copied

Ready to automate more?

Put your customer operations on auto-pilot

Ready to automate more?

Put your customer operations on auto-pilot

Ready to automate more?

Put your customer operations on auto-pilot