Data platform

In recent years, there has been a surge in demand for autonomous and scalable Data Platforms among companies. As the volume of data and the need for analysis requiring autonomy and freedom continue to grow, these platforms have transformed from optional tools to essential resources.

Today, Data Platforms often function as central repositories for all data, serving as workspaces for not only Data Analysts and Data Scientists but also Product Managers, Developers, Business Analysts, and anyone who understands the impact of data on all facets of business.

Although the market recognizes the benefits and requirements of Data Platforms, many companies still grapple with connecting disparate systems, scaling data, and providing sufficient tools to meet demands. Hindered by technological complexity, some companies suffer from painful failures due to unplanned decisions, while others are just beginning to build their own platforms. Overcoming these obstacles, however, could mark a turning point for businesses.

As a Technical Leader who has guided Data Platform teams across various sectors (Real Estate, Gaming, E-commerce), I will share the main challenges faced by these companies and the rationale behind building Data Platforms.

Challenges

The success of technology is gauged not only by its capabilities but also by the obstacles it can overcome. In this context, scalability and autonomy are the two most significant challenges of a Data Platform. To offer a practical perspective, I will outline three common issues I have encountered:

Rigid and non-scalable Data Warehouse solution

Data teams often find themselves tied to highly coupled Data Warehouse solutions. These solutions, which include poorly customizable Authentication and Authorization systems, limited automation Catalogs, and simplistic Workload Management, stem from well-known commercial systems and impede automation and customization.

Additionally, many commercial Data Warehouse solutions offer limited scalability options, lack transparency in scaling, and frequently necessitate considerable manual intervention.

Manual and centralized work in building

Analytical computations are increasingly consumed by decision-making applications or viewed by end-users seeking richer interpretation of their results. However, most market solutions fail to meet the Service Level Agreement (SLA) of applications or the volume of end-users.

In such cases, the most common approach involves manually building a system that applies various OLAP techniques to the massive data stored in Data Lakes before populating a database that meets the application SLA or massive end-user consumption.

These often homemade solutions generate numerous problems, such as the difficulty of reprocessing and managing data dispersed across multiple applications and databases. Most notably, this approach centralizes knowledge and autonomy within Data Engineering teams, obstructing the creative process of constructing Data Products.

Data teams centralizing all collection, transformation, and availability

This common issue frequently hampers the transformative potential of an autonomous data culture. A relic of the old Data team models, these teams centralized all knowledge of the business and data. The primary limitation is the linear scaling of Data Engineering teams with demand, resulting in constant bottlenecks.

While this centralizing model may still be effective for small or medium-sized companies or those requiring occasional data analysis, it is ill-suited for Exponential Organizations and rapidly growing companies, where the model reaches its limitations.

Solutions

With an understanding of the most common problems that prompt companies to seek autonomous and scalable solutions, I will outline the rationale behind the key decisions that guided me in building Data Platforms, as well as provide a simplified architecture overview. We will explore this process through two main topics: Consumption and Ingestion.