
How we built and integrated an AI developer: the EXANTE experience

My name is Alina, and I am Head of Web Applications at EXANTE. I am responsible for the technical strategy of our web projects, delivery efficiency and the development of engineering processes across teams. I also lead several initiatives, including our AI transformation.
When we started working with AI, we did not want to stop at individual tools for developers. We wanted to build an autonomous pipeline that could take on some tasks end to end. We also wanted to measure how far this approach improves developer productivity and shortens the time from task to finished result. This is how Codey came about. Codey is a multi-agent system of AI coding agents that independently covers much of the path from a Jira task to a working feature.
We began rolling it out six months ago. Since then, we have seen the following results:
- 250 tasks have gone through the system
- 75% of them were completed successfully
- Codey handles 7% of all tasks in the teams where it has been introduced
- It typically takes about 30 minutes from the start of coding to the deployment of a feature environment
A single powerful model is not enough to make AI part of a real development process. You also need high-quality requirements, well-prepared projects, measurable checks and clear rules for how the system operates.
From individual automations to an AI software engineer
We started with small automation tools: AI code review and analysis of affected modules based on the changed diff. They sped up individual stages after the code was written. However, they were not AI agents for software development that could complete the full cycle on their own, from analysing a Jira task and exploring the codebase to implementation and checks. I described this stage in my previous article.
That is why our next step was Codey. It is a system that explores the code, draws up a plan, writes the implementation and tests, checks the result and fixes any problems it finds, all on its own.
Codey currently works well with Node.js, React and Python projects. We began with local frontend and user interface (UI) tasks. We then moved on to full-stack changes that affect both frontend and backend.
Today we use the system across all projects in my area. These range from internal customer relationship management (CRM) and back-office systems to client-facing fintech services, including the client portal, the web trading terminal and the desktop application.
How Codey works: a multi-agent architecture
In architectural terms, Codey is an orchestration layer for AI agents that we developed in-house. By an AI agent, we mean a large language model (LLM) that works in a connected runtime environment, has a defined role and has access to the tools it needs. We use three runtimes: Claude Agent SDK, Codex and OpenCode. They allow agents to work with repositories, files and the terminal.

Codey connects them into a single agentic workflow. It assigns roles and stages, runs checks, returns work for revision and syncs the result with Jira, GitLab and continuous integration (CI).
How tasks reach Codey
The first scenario is built into sprint planning and depends on how far Codey has been introduced. At the start, the team manually selects well-described tasks and labels them for the system.
In teams with an established process, Codey and developers share a common backlog. By default, the system picks up suitable tasks. The team applies an AI skip label to tasks that a person should handle for specific reasons. Over time, we plan to extend this approach to all connected teams and projects.
The second scenario starts in Slack. The requester describes what they need, and Codey helps them navigate the system, clarifies the details and creates a Jira task. The system then completes the task, creates a merge request and returns to the same chat. It tags the requester, asks them to review the result and shares a link to the feature environment.
When Codey receives a task from Jira, it identifies the related GitLab repositories. Each project has predefined settings for:
- which repositories should receive the task
- which automated checks to run
- who to assign as reviewer
- how to deploy the feature environment and link the right frontend and backend versions
With these settings, Codey knows how to take a task through a specific project's process and prepare the result for the team's review.
AI agent roles in the system
Six core roles work on each task:
- Researcher studies the requirements and the codebase
- Planner draws up a plan
- Developer writes code and tests
- Reviewer checks the quality and correctness of changes
- QA (quality assurance) runs tests and technical checks
- Verifier compares the result with the mandatory requirements in the plan
The system uses different models for different roles. If Reviewer or the automated checks find a problem, Developer receives the task for rework.
After the result is handed over, the team reviews it. Codey takes into account comments on the code and business logic from the merge request and the Jira task. It then makes changes and runs the checks again.
To prevent a task from getting stuck in a loop of revisions, Codey makes no more than three iterations. If the comments are still unresolved, the task receives the AI failed status. The result is recorded in the statistics and feeds into the system's self-improvement loop. If needed, the run can be restarted manually.
Checks and error handling
Once a merge request is created, Codey runs the project's quality gates: tests, linters, security checks and AI code review. If the problem can be fixed, the task goes back for rework. If a check cannot run because of the project's configuration, Codey creates a Draft merge request and lists the checks that were not completed.
Most AI failed statuses have one of three causes: missing context, changes across several interdependent repositories and first tasks in a new project or technology stack.
Team review of the result
For each merge request, Codey identifies the affected modules. It notes in Jira which parts of the product need to be covered in regression testing. Once the checks pass, the frontend and backend are deployed to a separate feature environment, also known as a preview environment. There the requester can check the whole feature end to end.
If the result meets expectations, the requester clicks Confirm. The task then goes through team review and is included in the release build. Codey works with a human in the loop: product acceptance, the final decision and responsibility for the release stay with people.
Business challenges of adopting AI coding agents
Task quality. Like any member of a development team, Codey cannot complete a task well without clear requirements. That is why we introduced Jira task templates with a mandatory description of the acceptance flow and expected checks. These include requirements for unit and integration tests.
Team trust in Codey. At the start, we suggest beginning with routine tasks or refactoring, which go through specialist review in any case. As successful results accumulate, teams move from manual selection to a shared backlog for Codey and developers. They mark exceptions as AI skip. This builds trust gradually and gives the system more context.
A shift in workload to review and refinement. In the early stages of adoption, part of the workload moves to these stages. A developer needs to understand not only the task but also how Codey interpreted and implemented it. We reduce this workload with more precise task descriptions, detailed rules for each project and multi-level automated and human review.
Infrastructure costs. Separate backend feature environments increase infrastructure costs. To keep costs under control, we limit the lifetime of environments and use reduced test databases.
Technical challenges of a multi-agent system
System versatility. Projects differ in architecture, codebase structure, technology stack and development rules. You cannot give agents one set of instructions and expect equally high-quality results everywhere.
That is why we maintain a profile for each project. It describes the architecture, the codebase structure, the build and test commands and the local development rules. For recurring tasks, Codey has skills at three levels: global, language-specific and project-specific.
The system generates new versions of skills based on accumulated patterns. These are reviewed first and only then added to the pipeline. This approach to context engineering lets us adapt Codey's behaviour to each project.
Change quality. To achieve this, projects needed a simple, reproducible set-up, clear check commands and mature quality gates. Tests, linters, security checks and AI code review in the CI pipeline stop problematic changes before they add to the team's workload or reach production.
Choosing LLMs for different roles. To evaluate our coding agents, we run regular benchmarks. We look at how different combinations of models, roles and tools perform on the same set of general-purpose tasks. We then compare the solve rate. This is the share of solutions that passed the mandatory checks and received a score comparable to a human review. We also compare speed and cost of execution. For an independent evaluation, a result is never passed for review to a model from the same family.
The benchmarks showed that using the single most powerful model at every stage does not always give the best result. For some roles, a simpler model can be more effective than a complex one. It may follow the plan more closely, overcomplicate the solution less often and finish the task faster. That is why we regularly review how models are allocated across roles based on test results. Following the latest tests, we have already moved some roles to Codex models.
Continuous improvement. To improve Codey systematically, we use a Learning loop. It analyses logs and problematic sessions, identifies recurring errors and forms hypotheses. Each change is tested against the same benchmark and reaches production only after a measurable improvement.
In one of these cycles, we tested 20 hypotheses and kept nine. We also take feedback from developers and testers into account. In parallel, we optimise how the system works with code and context. RTK reduces the volume of command output, and an abstract syntax tree (AST) index speeds up the search for the right modules.
Next steps for Codey
In the near future, we will continue to develop Codey in several areas:
- Automated analysis of the shared backlog: the system will propose suitable tasks, and the team will exclude those that cannot be handed over to AI
- Developing the Learning loop: continuous analysis of human comments in merge requests and using this feedback to improve the engine automatically
- Wider coverage: connecting new requesters and departments, projects, technology stacks and systems to Codey
- Additional sources of context, such as documentation, Jira and work communications, and developing Codey as a foundation not only for development but also for analytical and product work on tasks
- A move to feature-based delivery, so that changes from Codey reach production faster after mandatory checks
Conclusion
Over six months, we have confirmed that an AI developer can be part of the working process. Connecting a powerful model to development is not enough. You need high-quality requirements, well-prepared projects, measurable checks and clear rules for how the system operates.
Our next step is to hand Codey the full cycle for some tasks, from task definition to production, while keeping the agreed rules and risk controls in place.
This article is provided to you for informational purposes only and should not be regarded as an offer or solicitation of an offer to buy or sell any investments or related services that may be referenced here. Trading financial instruments involves significant risk of loss and may not be suitable for all investors. Past performance is not a reliable indicator of future performance.



