Expanding a Retailer's Enterprise ETL Platform with New Data Flows

How do you keep adding new data feeds to a live enterprise ETL platform without putting business-critical loads at risk?

A large retail group's internal ETL team kept receiving requests for new data sources and deliveries, and each one had to reach production without disturbing the loads already running. Azati's engineers joined that team to build new data flows on an existing platform across three loader architectures, test them under load, and keep critical processes stable in production.

Ask us about your data platform
8 months

engagement length on a Time and Material basis

15-30

people in the client's internal ETL team that Azati's engineers worked within

3

loader architectures in daily work: RDBMS, Streaming, REST API

Technologies used

Apache NiFi
Apache NiFi
PostgreSQL
PostgreSQL
Hadoop HDFS
Hadoop HDFS
Hive
Hive
Greenplum
Greenplum
Oracle
Oracle
ClickHouse
ClickHouse
Apache Airflow
Apache Airflow
Apache Kafka
Apache Kafka
Python
Python

A working ETL platform with a growing queue of data requests

The client’s internal ETL team ran the platform that feeds analytics and key business processes for one of the largest retail chains in its market. Business teams, internal and external, kept asking for new sources and new data deliveries, and each request had to reach production without disturbing the loads already running.

When Azati joined, the platform was already in place: specialized loaders for different source types, including a modular loader for relational databases, a snapshot-based versioning mechanism, and a metadata configuration layer in PostgreSQL. Azati’s job was not to redesign it. The job was to build new flows on top of it, connect new source types, and keep critical loads stable in production.

Five demands shaped every new data flow on the platform

Every new flow had to meet five demands at once: fast delivery to the business, three different loader architectures, built-in history, stable processing at peak volumes, and production support that never paused.

Challenge 01

New data requests arrived faster than custom loading logic could keep up

The business regularly asked for new sources and additional data deliveries, and the loading process had to adapt to each one quickly:

  • Requests from both internal and external business customers
  • Analytical and operational use cases with different requirements
  • A need to configure new loads, not rewrite them for every request
#1
Challenge 02

Three loader architectures, one standard for reliability and data quality

Each loader was built for a specific source type and transfer method, but every flow had to meet the same requirements:

  • A modular loader for relational databases, with full, incremental, and repeat loads
  • Streaming processing for continuously arriving data, tied into the shared ETL infrastructure
  • REST API integration for service and external sources, under the same quality requirements
#2
Challenge 03

Every new source had to support historical analysis from the first day

The business needed data not only in its current state but also as it stood in earlier periods:

  • Business teams needed to restore data as of a past date
  • New flows had to use the platform’s existing versioning and change-storage mechanism correctly
  • A misconfigured flow would break historical analysis for that source
#3
Challenge 04

Large daily volumes left little room for a failed load

The platform processed large volumes from many corporate sources every day, so a failed load could directly affect analytics and business systems:

  • Stable processing of new flows under daily production load
  • Performance verified through load testing before a flow reached production
  • Lower risk of failures at peak data volumes across many sources
#4
Challenge 05

Development and production support shared the same engineers

Alongside new development, the team answered for the stability of every load already running:

  • Continuous monitoring of existing ETL processes
  • Fast incident resolution within first- and second-line support
  • On-call duty overnight, on weekends, and through holiday periods
#5

Why the client trusted Azati with production ETL work

Four practical reasons: experience with enterprise data platforms, a full cycle from analysis to support, speed from request to launch, and load testing before release rather than after an incident.

Experience with enterprise data platforms

Azati’s engineers have worked with large data volumes, distributed storage systems, and production ETL processes, which is what complex corporate data environments demand from day one.

A full cycle, from system analysis to production

Azati covered the whole path of a data flow: system analysis and task definition, development, testing including load testing of the ETL cluster, and support once the flow was live.

Speed from request to launch

Working inside the client’s own team let Azati’s engineers pick up new initiatives quickly, adapt to changing requirements, and shorten the time between a business request and a running flow.

Load testing before release, not after an incident

For critical changes, Azati ran load tests and performance assessments in advance, so the behavior of a new flow under peak volumes was known before release.

Is every new data feed a custom project for your ETL team?

If each request to connect a source turns into new code instead of new configuration, the bottleneck is usually the process around the platform, not the team. Tell us what your pipeline takes too long to add.

Ask us about your data platform

How Azati built new data flows across six areas of the platform

Azati’s work on the platform ran along six lines: new flows on the relational loader, streaming, historical storage, metadata-driven configuration, load testing, and access-controlled production support. The architecture belonged to the client. The flows, tests, and support were where Azati’s engineers put in the hours.

01

New data flows on the relational database loader

Azati’s engineers created and connected new data flows on the platform’s existing modular loader for relational databases, inside an ETL layer built on Apache NiFi, Apache Airflow, and Apache Sqoop. Each flow went from source mapping to full, incremental, or repeat loading, and stayed under Azati’s support after launch.

Key capabilities:
  • Connecting new sources and target tables
  • Full, incremental, and repeat load scenarios
  • Handling source schemas and their changes
  • Flow support after launch
02

Streaming data processing

For sources that deliver data continuously, Azati designed and configured real-time processing flows and connected them to the shared ETL infrastructure, so streaming data met the same reliability and quality standards as batch loads.

Key capabilities:
  • Ingesting and processing streaming data
  • Flows built for specific business scenarios
  • Integration with the shared ETL infrastructure
03

Historical data storage

The platform stores periodic reference snapshots and the changes between them, so any table can be restored to its state on a chosen date. Azati didn’t build this mechanism. Azati made sure every new source used it correctly, so a new flow supported historical analysis from the day it went live.

Key capabilities:
  • Periodic reference snapshots
  • Storage of changes between snapshots
  • Restoring a table’s state for a chosen date
  • Snapshot frequency set through configuration
04

Metadata-driven configuration

Load parameters, processing rules, and source settings live in a separate configuration layer in PostgreSQL. That let Azati’s engineers connect new tables and adapt loads to new requirements by changing configuration, not code.

Key capabilities:
  • Load configuration stored in PostgreSQL
  • Table fragmentation and parallelization settings
  • Incremental loading rules
  • Versioned column mapping schemas
05

Load testing of ETL processes

Before a new flow went to release, Azati tested loader performance under the client’s existing testing procedure: modeled load, resource usage analysis, and processing time for large data volumes.

Key capabilities:
  • Processing time for a set data volume
  • Loader throughput measurement
  • Controlled test load generation
  • CPU, RAM, and queue state tracking
06

Access control and platform support

The platform separates the rights of developers, service processes, and support specialists. Within that model, Azati handled both the development of new flows and their support in production, including on-call monitoring and defect analysis.

Key capabilities:
  • Separate development and operations rights
  • Service roles for loading and monitoring
  • View-only role for support specialists
  • Dedicated rights for service operations

Screenshots

Expanding a Retailer's Enterprise ETL Platform with New Data Flows

What Azati delivered

AreaAzati contribution
RDBMS data flowsCreated and connected new data flows on the existing modular relational database loader
StreamingDesigned and configured flows for sources that deliver data in real time
Historical storageApplied the existing versioning and change-storage mechanism to every new flow
ETL configurationUsed the PostgreSQL metadata layer to manage sources and processing parameters
Performance controlLoad tested and assessed the performance of new flows under the existing procedure
Platform supportTook part in production support, loader monitoring, and incident resolution

Three loader architectures, one delivery standard

Loader architectureSource typeAzati’s work
RDBMSRelational databasesNew flows on the existing modular loader with full, incremental, and repeat loads
StreamingContinuously arriving dataReal-time processing flows connected to the shared ETL infrastructure
REST APIService and external systemsDevelopment and support of API-based integration flows under the same reliability requirements

Security

The platform ran inside the client’s own internal infrastructure. Access followed a role model that separates developer, service-process, and support rights, and Azati’s engineers worked under the client’s operational and access policies throughout the engagement.

Team composition

Azati’s ETL engineers worked embedded inside a client-led internal ETL team of 15 to 30 people.

  • ETL Developers responsible for system analysis, new data flow development, load testing, and first- and second-line production support, working on a Time and Material basis from October 2023 to July 2024.

How the engagement was delivered

Time and Material

Azati’s engineers worked under a T&M model for 8 months, embedded directly inside the client’s internal ETL team rather than delivering as a separate external unit.

Mixed methodology across staff and contracted engineers

Most of the client’s own staff worked in Agile sprints, while engineers brought in for specific business requests, Azati’s team included, worked in Kanban, picking up and completing individual flow requests as they came in rather than committing to fixed sprint scope.

More data flows in production, fewer load risks, faster source onboarding

The engagement produced four results the business felt in daily work: a steady stream of new flows, more reliable loads, faster onboarding of new sources, and stable critical integrations.

More data flows reached production

New sources and ETL flows were connected on a regular basis at the request of internal and external business customers.

Loads became more reliable

Following the platform’s architectural standards and performance control procedures reduced the risks of processing large data volumes.

New sources went live faster

Ready-made configuration mechanisms let new ETL flows launch without changing the platform’s core logic.

Critical integrations stayed stable

Production support kept business-critical integration flows running reliably.

Strategic wins

Three advantages outlasted the engagement: a sounder base for analytics, development shaped by production, and ETL that grows through configuration rather than code.

A reliable base for analytics and decisions

Centralized data processing and support for historical slices improved the quality of data available to business teams. A report is only as trustworthy as the load behind it.

Development shaped by production

Azati’s part in production support meant new flows were built with the demands of live operation in mind. Engineers who answer for a failed load at night build the next flow differently.

Scaling ETL through configuration, not code

Configuration automation and performance control gave the platform room to grow with data volumes, and turned each new request into one more configured flow instead of one more custom loader.

The described expertise is relevant for

This engagement model is unlikely to be the right fit for

  • Teams that want a fixed, one-time integration project rather than ongoing embedded capacity inside their own data function.
  • Organizations whose production support cannot be shared with an external engineering partner.
  • Companies looking for a data platform designed from scratch rather than development on an existing one.
  • Environments with a handful of stable sources and few new data requests.

Frequently asked questions

Azati’s engineers built new data flows on an existing enterprise ETL platform across relational, streaming, and REST API loaders. The work also covered system analysis, load testing, historical data setup, and first- and second-line production support over 8 months.

No, the platform and its loader architecture were already in place before Azati joined the client’s ETL team. Azati’s role was to develop new flows on that architecture, connect new source types, and keep critical loads stable in production.

New flows use the platform’s existing snapshot and change-storage mechanism, so a table can be restored to any chosen date. Azati configured each new source to use this mechanism from launch, which meant historical analysis worked from the first day.

Load parameters, processing rules, and source settings live in PostgreSQL, so new tables connect through configuration instead of code changes. That let Azati’s engineers adapt loads to new business requirements faster and without touching the platform’s core logic.

Every new flow was load tested under the client’s existing procedure before release, with modeled load and resource tracking. Azati measured processing time for set data volumes, loader throughput, and CPU, RAM, and queue state to catch problems before production.

The on-call model meant monitoring loads overnight, on weekends, and through holiday periods, not only during business hours. Azati’s engineers carried first- and second-line support alongside development, including incident resolution and defect analysis for critical ETL processes.

Azati worked on a Time and Material basis, embedded inside a client-led internal ETL team of 15 to 30 people. The role was hands-on flow development, testing, and production support within that team, not ownership of the whole platform.

Last updated

Got a job for Azati? Let’s talk business!

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

What's next?

  • 1. Tell Us Your Story
    Describe your project. We come back within 24 hours with team availability and a rough plan. NDA on request before the first call.
  • 2. Get Your Roadmap
    Receive a detailed proposal with scope, team composition, timeline, and costs tailored to your goals.
  • 3. Start Building
    Azati aligns on details, finalize terms, and launch your project with full transparency.