People sitting around a table

adesso Blog

One specific transformation project, two different challenges

The starting point is a large-scale modernisation project in the financial sector. Existing ETL processes, data assets and business rules are to be migrated from a system landscape that has evolved over many years to a modern lakehouse architecture. At the same time, new data products are being developed on the new Databricks environment.

The demands are high: the project team must understand complex legacy logic, reliably transfer business rules and implement new transformation pipelines. In doing so, neither data quality nor security, traceability or regulatory compliance must be compromised.

The project uses Agentic AI for these tasks. This significantly reduces the analysis and implementation effort that would take several weeks using a largely manual approach. However, it quickly became clear during the project that a single agent harness is not equally suitable for all tasks.

The key factor is, above all, how much context, access and freedom of action an agent requires for the task in question. An agent harness refers to the technical environment in which an AI agent can access information and tools and carry out tasks with a defined degree of autonomy.

VS Code and GitHub Copilot for reverse engineering

A key challenge in the project is the reverse engineering of the existing ETL processes. The workflows to be analysed have evolved over many years and, in some cases, are only incompletely documented. Relevant logic is found, amongst other places, in source code, stored procedures, configuration files, various repositories and technical documentation.

A harness based on VS Code and GitHub Copilot is used for this task. It enables the engineering agent to work beyond the confines of individual files and systems. Depending on the tools and permissions granted, the agent can:

  • examine and compare code across multiple repositories,
  • track dependencies between components and data flows,
  • analyse SQL, scripts and configuration files,
  • execute shell commands and tests,
  • document transformation logic and business rules,
  • prepare the migration of legacy logic to the Lakehouse architecture.

This breadth is particularly important in this specific project because the relevant information is not available in a single development environment or platform. The agent must be able to identify relationships across system and repository boundaries.

The estimated effort for reverse engineering, using a largely manual approach, was around 500 person-days. Thanks to the agent’s support, the actual effort was reduced to fewer than 100 person-days. The results demonstrate the considerable potential of this approach.

Greater flexibility also means greater risk

The flexibility required for reverse engineering is, at the same time, the greatest governance challenge posed by the VS Code-based harness. Developers can install extensions, integrate different models or tools, and set up their own workflows. Without central guidelines, this can lead to individual agent environments that are difficult to control – effectively creating the agent-based equivalent of shadow IT.

This is particularly critical in a project within the financial sector. Unauthorised services could process sensitive code, metadata or data. Agents could operate with overly broad permissions or make changes that are not adequately logged and verified.

Clear guidelines are therefore essential for project deployment. These include, in particular:

  • centrally approved models, extensions and tools,
  • role-based access and authorisation concepts,
  • traceable logging of agent activities,
  • mandatory testing and review processes,
  • defined human approval points,
  • reusable and centrally maintained configurations.

Experience from the project shows that the greater an agent’s scope for action, the more robust the technical and organisational control mechanisms must be.

Databricks Genie Code for implementation in the Lakehouse

In parallel with the reverse engineering of the existing ETL processes, the project team is undertaking a second key task: the development of new data products within the Databricks environment. To this end, it creates dbt-Core models and implements the necessary transformations in line with the Medallion architecture.

Databricks Genie Code is utilised in this phase. Unlike the more broadly scoped VS Code Harness, Genie Code operates more closely within the shared platform context. It can assist engineers in:

  • understanding existing data sets and Databricks artefacts,
  • generate and optimise SQL and Python code,
  • develop and refine dbt Core models,
  • implement transformations across the Medallion architecture (Bronze, Silver, Gold),
  • work with notebooks and pipelines,
  • and deliver new data products and use cases more quickly.

The narrower platform focus may limit tasks outside of Databricks. Where extensive interaction with external repositories, infrastructure tools or heterogeneous legacy systems is required, Genie Code may therefore be less suitable.

Within the target environment, however, this limitation becomes an advantage for the project. The team works in a more standardised context with shared data sets, permissions, development artefacts and governance guidelines. This makes it easier to establish and scale consistent engineering practices.

Genie Code also supports database administrators within the project who have previously worked primarily with Oracle and PL/SQL. As they transition to Databricks and dbt Core, the agent-based support helps them to understand and apply new technologies, syntax and development patterns more quickly.

Broad understanding, standardised implementation

The project demonstrates that the two harnesses should not be viewed as competing solutions. They fulfil different roles within the same transformation:

VS Code with GitHub Copilot provides the necessary cross-system access for reverse engineering the existing ETL landscape. This flexibility enables a comprehensive analysis but requires strict governance.

Databricks Genie Code supports the subsequent implementation within the new Lakehouse environment. The clearer platform focus promotes standardisation, shared development patterns and controlled workflows.

The combined use brings together the necessary breadth in reverse engineering with a more standardised implementation on the target architecture.

The dividing line does not run strictly between reverse engineering and implementation. Both harnesses can, in principle, support engineering tasks. What is decisive, rather, is their respective operational context: reverse engineering requires access across numerous technical boundaries, whilst implementation benefits from a shared and controlled platform context.

Productivity requires verifiable quality

The productivity gains achieved in the project are substantial. However, speed alone is not a sufficient measure of success. Agent-generated code and implementations must meet the same quality requirements as manually created outputs.

These include:

  • automated tests for transformation and business logic,
  • peer reviews of generated or modified code,
  • checks on data quality and data provenance,
  • security and authorisation controls,
  • traceable documentation,
  • clearly assigned human responsibility.

In this project, therefore, Agentic AI does not replace the expertise of data engineers, architects and business stakeholders. Above all, it reduces time-consuming analysis and implementation work, thereby creating more scope for validation, technical clarification and architectural decisions.

Conclusion

This specific modernisation project in the financial sector makes it clear: the most effective agent harness is the one that suits the specific task at hand.

VS Code with GitHub Copilot offers the scope required for reverse-engineering a heterogeneous and historically evolved ETL landscape. Databricks Genie Code, on the other hand, provides a more standardised framework for implementation with dbt Core within the Lakehouse architecture.

Combining both approaches can significantly accelerate analysis and implementation. This requires that flexibility, security and governance be considered together from the outset. Particularly in a regulated environment, sustainable added value is not created by maximising autonomy, but rather by defining a deliberate scope of action for each agent.

Would you like to design modern data platforms and deploy generative and agentic AI in a secure and value-driven manner? adesso can support you in this. Further information and relevant content can be found on the adesso website under Generative AI.

Picture Mathias  Weber

Author Mathias Weber

Mathias Weber is a data architect and managing consultant in the Data Platform Solutions business unit within the Data & AI business line at adesso. With a technology-neutral approach to enterprise-wide data and AI platforms, he helps organisations develop scalable and sustainable solutions that combine technological possibilities with strategic business requirements.