Arthur Reis Data & AI Engineer
Microsoft Fabric · Databricks · Azure
Data and AI Engineer specialized in Microsoft Fabric, Databricks and Azure. I build modern data platforms, lakehouse architectures, high-performance pipelines and data quality frameworks, and explore applied AI with multi-agent systems. I share what I learn through talks, YouTube videos, blog posts, LinkedIn and open source projects.
Who is Arthur
Data and AI Engineer specialized in Microsoft Fabric, Databricks and Azure. I build modern data platforms, lakehouse architectures, high-performance pipelines and data quality frameworks, and explore applied AI with multi-agent systems. I share what I learn through talks, YouTube videos, blog posts, LinkedIn and open source projects.
- Data Engineering
- SQL & Spark
- Quality & Governance
- AI & Agents
Tools & technologies
Data Engineering
SQL & Spark
Quality & Governance
AI & Agents
Experience
- Sep 2025 — Present
Data Engineer · Power Tuning
Brazil · RemoteNational and international data platform strategy and modernization with Microsoft Fabric and Databricks.
- Designed and deployed a Data Platform for a major Brazilian multinational using Microsoft Fabric and Databricks, aligning technical solutions with business and governance objectives.
- Architected and deployed data workflows via Infrastructure as Code (IaC) through Azure Data Factory, ensuring scalable, maintainable and auditable pipelines.
- Designed and deployed an entire Lakehouse and AI products using Databricks and Azure.
- Implemented Databricks Genie Spaces as a conversational analytics interface over the platform data.
Microsoft Fabric Databricks Azure Azure Data Factory PySpark SQL Python - May 2025 — Sep 2025
Data Engineer · Kumulus
Brazil · RemoteData strategy, platform design and implementation for national and international clients.
- Led data platform strategy and Lakehouse architecture design for a U.S.-based multinational, defining governance standards, data models and platform structure aligned with client goals.
- Contributed to data platform modernization for a U.S. government organization, mapping data maturity and driving adoption with international stakeholders.
- Designed and implemented Lakehouse solutions for two Brazilian companies, including governance frameworks and data quality controls.
- Mentored and supervised junior Data Engineers, contributing to team topology definition.
- Deployed Microsoft Copilot to increase data team productivity across clients.
Databricks Azure Python SQL Delta Lake - Apr 2024 — May 2025
Data Engineer · Paytime
Brazil · RemoteOwned the company's end-to-end data strategy from the ground up: infrastructure, governance, quality and analytics.
- Designed and implemented the Data Lakehouse (Databricks, Spark, SQL, MongoDB, AWS) as the foundation for scalable data products and self-service analytics.
- Established data governance practices including quality testing pipelines, documentation standards and access control policies.
- Designed and delivered the company's first data products: interactive Power BI dashboards for business stakeholders.
- Implemented AWS data services (S3, Lambda, Glue, IAM) for a secure and scalable data infrastructure.
- Explored and implemented Databricks Genie Spaces for conversational analytics over company data.
Databricks Apache Spark AWS MongoDB Power BI Python SQL - Sep 2023 — Mar 2024
Data Analyst · Piwi
Brazil · RemoteManaged the company's end-to-end data strategy, from ingestion to analytics enablement.
- Led the development of the company's Lakehouse on GCP, defining the data architecture strategy and enabling scalable data management.
- Designed and launched the company's first analytics data products using Looker Studio, enabling self-service BI.
- Developed data extraction and cleaning pipelines using Python, SQL and MongoDB.
GCP Python SQL MongoDB Looker Studio - Mar 2022 — Jan 2023
Data Analyst · Secretaria Estadual de Educação do ES
Espírito Santo · BrazilSupporting data-driven decision-making for public educational policy in Espírito Santo.
- Conducted data analysis and strategic reporting on Indigenous, Quilombola and Rural schools to inform educational policy decisions.
- Produced educational resources and led training sessions on data practices for school staff.
Power BI Python SQL Excel - Mar 2018 — Nov 2022
Researcher / Data Analyst · UFES — Universidade Federal do Espírito Santo
Vitória, ES · BrazilMaster's and PhD research focused on historical data analysis and social network modeling.
- Performed ETL on historical documents and built data pipelines using Python.
- Conducted social network analysis (Twitter) through API integration using Python.
- Built visualizations and dashboards in Power BI to communicate research findings to academic stakeholders.
- Developed the website www.jornaisdaindependencia.com.br.
Python R Power BI SQL Excel
Featured projects
Professional and personal projects
data-agents-copilot
Automatic dispatch system that routes data tasks (SQL, PySpark, pipelines, governance) to 15 specialized AI agents. Runs via CLI, Chainlit web, or directly in VS Code Chat.
Fork of the original data-agents project by Thomaz Rossito — adapted to run with GitHub Copilot Chat API, adding automatic naming governance, collaborative multi-agent workflows, structured Knowledge Base, episodic memory system and peer-to-peer QA protocol.
Databricks Platform Refactoring & Optimization
Refactoring and optimization project for a Databricks-based data platform focused on performance, maintainability, and critical SLA compliance. After a complete framework redesign and optimizations in PySpark, SQL, and platform architecture, the solution successfully achieved the 15-minute SLA while becoming more scalable, stable, and easier to maintain.
ETL Optimization on Databricks
Optimization of critical ETL pipelines on Databricks using BroadcastJoin, CLUSTER BY, OPTIMIZE and incremental reads. Pipelines achieved up to 94% runtime reduction, with direct impact on DBU consumption and processing windows.
Lakehouse on Azure
Designed and implemented a corporate Lakehouse for a major Brazilian company using the medallion architecture (Bronze, Silver and Gold), integrating multiple data sources with focus on ERP Protheus. Business-domain workspaces integrated into a central Engineering workspace for governance and standardization. Data loads orchestrated via Azure Data Factory with parameterized, SLA-controlled pipelines.
App Modernization
Full migration from an On-Premise Data Warehouse to Lakehouse architecture on Microsoft Fabric: dimensional modeling, ingestion and transformation pipelines, SQL procedure migration to Lakehouse, and publishing reports and dashboards in Power BI connected to the Lakehouse.
Stack: Microsoft Fabric, OneLake, Data Factory, PySpark, Synapse Data Warehouse, Power BI, T-SQL, Delta Lake.
AI Implementation and Testing
Integration of company data with Databricks AI (Genie) to accelerate processes and foster a Data-Driven culture. Conducted testing and promoted tool adoption across the company.
MongoDB to Databricks Data Migration
Migration of a MongoDB database to Databricks using MongoDB Data Federation, AWS S3, and incremental load routines with Spark. The result was a regular update routine for the company's most critical data, including transaction tables with millions of rows.
Lakehouse Development
Implementation of a Lakehouse on Databricks integrating all data from a Fintech using SQL, Spark, Databricks and AWS. The result was improved performance in analytics and report/dashboard updates.
Dashboard Portal
Created a Dashboard Portal that reduced Power BI license costs from R$45/license to R$5/license. The cost reduction and platform capabilities led to extending the Portal to over 500 company clients, turning it into a data product.
Gender Classification using AI
Automatic gender classification in a legacy database using Python's gender_guesser.detector library. Achieved approximately 87% accuracy (calculated from a reduced corpus).
Repositories
Public code on GitHub
data-agents-copilot
Multi-agent system for data engineering with 15 specialist agents, episodic memory and MCP servers.
dqx_framework
Data quality framework built on Databricks DQX.
fabric-documenter
Automatic documentation of Microsoft Fabric notebooks and DAX measures via API.
great_expectations_framework
Data validation and contracts with Great Expectations.
auto_comments_columns_databricks
Automatic generation of table column descriptions in Databricks.
snowflake_cortex
Exploring GenAI and data analytics with Snowflake Cortex.
Talks & events
Data Quality on Databricks
Data quality frameworks, data contracts, metrics and anti-patterns — ISO-25012 and DAMA DMBOK applied to the medallion architecture.
Modern data architectures: Snowflake, Databricks and a bit of AI
Comparison of modern data architectures — Snowflake and Databricks — and how AI is transforming the way data platforms are built and operated.
Certifications
Databricks Certified Data Engineer Associate
Databricks VerifyApache Spark Developer
Databricks VerifyDatabricks Certified Data Analyst Associate
Databricks VerifyDP-750: Microsoft Fabric Analytics Engineer
Microsoft VerifyDP-700: Microsoft Fabric Data Engineer
Microsoft VerifyDP-900: Azure Data Fundamentals
Microsoft VerifyAZ-900: Azure Fundamentals
Microsoft VerifyAI-900: Microsoft Azure AI Fundamentals
Microsoft VerifyGCP Associate Cloud Engineer
Google VerifyApache Airflow
Astronomer VerifyMongoDB for SQL Professionals
MongoDBGoogle Data Analytics
Google / Coursera VerifyGoogle Project Management
Google / Coursera VerifyLet's talk
Open to projects, talks and technical exchanges.