A Arthur Reis

Arthur Reis Data & AI Engineer

Microsoft Fabric · Databricks · Azure

Data and AI Engineer specialized in Microsoft Fabric, Databricks and Azure. I build modern data platforms, lakehouse architectures, high-performance pipelines and data quality frameworks, and explore applied AI with multi-agent systems. I share what I learn through talks, YouTube videos, blog posts, LinkedIn and open source projects.

Arthur Reis
5+ Years of experience
13 Certifications
10 Core projects
2 Talks
About

Who is Arthur

Data and AI Engineer specialized in Microsoft Fabric, Databricks and Azure. I build modern data platforms, lakehouse architectures, high-performance pipelines and data quality frameworks, and explore applied AI with multi-agent systems. I share what I learn through talks, YouTube videos, blog posts, LinkedIn and open source projects.

  • Data Engineering
  • SQL & Spark
  • Quality & Governance
  • AI & Agents
Tech Stack

Tools & technologies

Data Engineering

Microsoft Fabric Databricks Azure PySpark Delta Lake Azure Data Factory AWS GCP Snowflake

SQL & Spark

SQL Spark SQL Python DAX R

Quality & Governance

Unity Catalog DQX Great Expectations Data Contracts Power BI

AI & Agents

Multi-agent Systems RAG Prompt Engineering Mosaic AI MCP
Journey

Experience

  1. Sep 2025 — Present

    Data Engineer · Power Tuning

    Brazil · Remote

    National and international data platform strategy and modernization with Microsoft Fabric and Databricks.

    • Designed and deployed a Data Platform for a major Brazilian multinational using Microsoft Fabric and Databricks, aligning technical solutions with business and governance objectives.
    • Architected and deployed data workflows via Infrastructure as Code (IaC) through Azure Data Factory, ensuring scalable, maintainable and auditable pipelines.
    • Designed and deployed an entire Lakehouse and AI products using Databricks and Azure.
    • Implemented Databricks Genie Spaces as a conversational analytics interface over the platform data.
    Microsoft Fabric Databricks Azure Azure Data Factory PySpark SQL Python
  2. May 2025 — Sep 2025

    Data Engineer · Kumulus

    Brazil · Remote

    Data strategy, platform design and implementation for national and international clients.

    • Led data platform strategy and Lakehouse architecture design for a U.S.-based multinational, defining governance standards, data models and platform structure aligned with client goals.
    • Contributed to data platform modernization for a U.S. government organization, mapping data maturity and driving adoption with international stakeholders.
    • Designed and implemented Lakehouse solutions for two Brazilian companies, including governance frameworks and data quality controls.
    • Mentored and supervised junior Data Engineers, contributing to team topology definition.
    • Deployed Microsoft Copilot to increase data team productivity across clients.
    Databricks Azure Python SQL Delta Lake
  3. Apr 2024 — May 2025

    Data Engineer · Paytime

    Brazil · Remote

    Owned the company's end-to-end data strategy from the ground up: infrastructure, governance, quality and analytics.

    • Designed and implemented the Data Lakehouse (Databricks, Spark, SQL, MongoDB, AWS) as the foundation for scalable data products and self-service analytics.
    • Established data governance practices including quality testing pipelines, documentation standards and access control policies.
    • Designed and delivered the company's first data products: interactive Power BI dashboards for business stakeholders.
    • Implemented AWS data services (S3, Lambda, Glue, IAM) for a secure and scalable data infrastructure.
    • Explored and implemented Databricks Genie Spaces for conversational analytics over company data.
    Databricks Apache Spark AWS MongoDB Power BI Python SQL
  4. Sep 2023 — Mar 2024

    Data Analyst · Piwi

    Brazil · Remote

    Managed the company's end-to-end data strategy, from ingestion to analytics enablement.

    • Led the development of the company's Lakehouse on GCP, defining the data architecture strategy and enabling scalable data management.
    • Designed and launched the company's first analytics data products using Looker Studio, enabling self-service BI.
    • Developed data extraction and cleaning pipelines using Python, SQL and MongoDB.
    GCP Python SQL MongoDB Looker Studio
  5. Mar 2022 — Jan 2023

    Data Analyst · Secretaria Estadual de Educação do ES

    Espírito Santo · Brazil

    Supporting data-driven decision-making for public educational policy in Espírito Santo.

    • Conducted data analysis and strategic reporting on Indigenous, Quilombola and Rural schools to inform educational policy decisions.
    • Produced educational resources and led training sessions on data practices for school staff.
    Power BI Python SQL Excel
  6. Mar 2018 — Nov 2022

    Researcher / Data Analyst · UFES — Universidade Federal do Espírito Santo

    Vitória, ES · Brazil

    Master's and PhD research focused on historical data analysis and social network modeling.

    • Performed ETL on historical documents and built data pipelines using Python.
    • Conducted social network analysis (Twitter) through API integration using Python.
    • Built visualizations and dashboards in Power BI to communicate research findings to academic stakeholders.
    • Developed the website www.jornaisdaindependencia.com.br.
    Python R Power BI SQL Excel
Projects

Featured projects

Professional and personal projects

Feb 2026 — Present

data-agents-copilot

Automatic dispatch system that routes data tasks (SQL, PySpark, pipelines, governance) to 15 specialized AI agents. Runs via CLI, Chainlit web, or directly in VS Code Chat.

Fork of the original data-agents project by Thomaz Rossito — adapted to run with GitHub Copilot Chat API, adding automatic naming governance, collaborative multi-agent workflows, structured Knowledge Base, episodic memory system and peer-to-peer QA protocol.

Python Multi-agent LLM Databricks Fabric
Jan 2026 — May 2026 Power Tuning

Databricks Platform Refactoring & Optimization

Refactoring and optimization project for a Databricks-based data platform focused on performance, maintainability, and critical SLA compliance. After a complete framework redesign and optimizations in PySpark, SQL, and platform architecture, the solution successfully achieved the 15-minute SLA while becoming more scalable, stable, and easier to maintain.

Databricks PySpark SQL Delta Lake
Jul 2025 — Sep 2025

ETL Optimization on Databricks

Optimization of critical ETL pipelines on Databricks using BroadcastJoin, CLUSTER BY, OPTIMIZE and incremental reads. Pipelines achieved up to 94% runtime reduction, with direct impact on DBU consumption and processing windows.

Databricks PySpark Spark SQL Delta Lake
Sep 2025 — Jan 2026 Power Tuning

Lakehouse on Azure

Designed and implemented a corporate Lakehouse for a major Brazilian company using the medallion architecture (Bronze, Silver and Gold), integrating multiple data sources with focus on ERP Protheus. Business-domain workspaces integrated into a central Engineering workspace for governance and standardization. Data loads orchestrated via Azure Data Factory with parameterized, SLA-controlled pipelines.

Microsoft Fabric Azure Data Factory Delta Lake Medallion
May 2025 — Aug 2025 Kumulus

App Modernization

Full migration from an On-Premise Data Warehouse to Lakehouse architecture on Microsoft Fabric: dimensional modeling, ingestion and transformation pipelines, SQL procedure migration to Lakehouse, and publishing reports and dashboards in Power BI connected to the Lakehouse.

Stack: Microsoft Fabric, OneLake, Data Factory, PySpark, Synapse Data Warehouse, Power BI, T-SQL, Delta Lake.

Microsoft Fabric PySpark Power BI T-SQL
Feb 2025 — Jul 2025 Paytime

AI Implementation and Testing

Integration of company data with Databricks AI (Genie) to accelerate processes and foster a Data-Driven culture. Conducted testing and promoted tool adoption across the company.

Databricks Genie AI Analytics
Mar 2025 — Mar 2025 Paytime

MongoDB to Databricks Data Migration

Migration of a MongoDB database to Databricks using MongoDB Data Federation, AWS S3, and incremental load routines with Spark. The result was a regular update routine for the company's most critical data, including transaction tables with millions of rows.

MongoDB Databricks Spark AWS S3
Nov 2024 — Mar 2025 Paytime

Lakehouse Development

Implementation of a Lakehouse on Databricks integrating all data from a Fintech using SQL, Spark, Databricks and AWS. The result was improved performance in analytics and report/dashboard updates.

Databricks Spark SQL AWS Delta Lake
Oct 2024 — Jan 2025 Paytime

Dashboard Portal

Created a Dashboard Portal that reduced Power BI license costs from R$45/license to R$5/license. The cost reduction and platform capabilities led to extending the Portal to over 500 company clients, turning it into a data product.

Power BI Databricks SQL
Nov 2023 — Dec 2023 Piwi

Gender Classification using AI

Automatic gender classification in a legacy database using Python's gender_guesser.detector library. Achieved approximately 87% accuracy (calculated from a reduced corpus).

Python AI Data Quality
Blog

Articles on Medium

Writing about data engineering and AI

Community

Talks & events

SQL Saturday #1139 Joinville

Data Quality on Databricks

April 11, 2026Joinville, SC · Brazil

Data quality frameworks, data contracts, metrics and anti-patterns — ISO-25012 and DAMA DMBOK applied to the medallion architecture.

SQL Saturday #1129 Vitória

Modern data architectures: Snowflake, Databricks and a bit of AI

December 06, 2025Vitória, ES · Brazil

Comparison of modern data architectures — Snowflake and Databricks — and how AI is transforming the way data platforms are built and operated.

Contact

Let's talk

Open to projects, talks and technical exchanges.