About

I’ve been a data engineer for over 12 years — 6 of them in software and web development, and the last 6 focused on data engineering, streaming platforms, and data infrastructure at scale.

I’m currently a Senior Staff Software Engineer at will bank. Before that, I led the data platform at MonetizeMore as Lead Data Engineer, and held data platform Tech Lead roles at Escale Digital and Wildlife Studios. Earlier in my career I worked as a Data Engineer at Grupo ZAP and VAGAS.com — always on the side of building the infrastructure that supports product and business decisions, not just the final reports.

How I got into data

I started out as a backend developer, spending a good chunk of that time in the Ruby on Rails ecosystem — through software boutiques and consultancies (Codeminer42, Locaweb) and a high-traffic e-commerce project at Walmart.com Brazil, where I learned a lot about operating critical systems under real load. Even earlier, during my Computer Science degree at Universidade Estadual de Londrina, I spent time doing research applying Machine Learning to biological and medical problems — that’s probably where my curiosity about data first started competing with my love of building software.

The move into data engineering happened naturally: as the systems I built grew, the more interesting problems stopped being “how do I serve this request” and became “how do I move, transform, and trust data at the right scale.”

What I do

I mostly work on moving traditional ETL architectures toward agile, streaming-first ones — SQL on Hadoop (Athena/Presto), distributed processing with Spark, orchestration with Airflow, and streaming platforms built on Kafka. I also have experience with distributed transactional databases (Cassandra, HBase) and metadata governance tooling (Atlas, Hive Metastore).

A good part of my work over the last few years hasn’t been just code — it’s also technical leadership: guiding teams through architecture migrations, setting engineering standards, and helping teams move away from a centralized model (one person or team holding all the knowledge) toward a data platform that any team can use with autonomy.

Writing isn’t new to me either — I’ve previously published on data science and machine learning topics, so this blog is less of an experiment and more a continuation of a habit.

About this blog

Here I write about the things that occupy most of my day-to-day thinking: data platform execution models, streaming, lakehouses, and the engineering decisions behind them — no big promises, just what I’ve learned trying to solve real problems, usually after getting it wrong the first time.

You can find me on GitHub and LinkedIn.

ESC

Type to search...