Initial Phase of Implementing Airbyte

Context

This blog post is about the thought process behind our first steps towards implementing Airbyte. We have yet to implement it, as we are still exploring whether Airbyte is a good fit within our tech stack.

The Discovery Process

One of the tasks assigned to me was writing a ticket regarding the implementation of Airbyte as part of the company’s data ingestion pipeline. Initially, this task raised many questions in my mind. What exactly is Airbyte? Why does the company want to use it? Why not continue using the current pipelines? What does the current pipeline even look like?

My Research & Discovery Process
📋
STEP 1
Assigned Task
Receive a task I know very little about
➔
❓
STEP 2
Ask Questions
Identify gaps and clarify the context
➔
🔍
STEP 3
Research Independently
Demo videos, docs, internal slides
➔
💡
STEP 4
Build Understanding
Piece together the full picture
Figure 1 — From confusion to clarity: my process when faced with an unfamiliar task.

The Data Ingestion Pipeline

Rather than immediately jumping into writing the ticket, I realised that I first needed to properly understand the problem and the surrounding context. To answer these questions, I began researching independently. I watched demo videos to understand how Airbyte works and explored the different connectors available on the platform. At the same time, I went through Vatico’s internal slides and diagrams to better understand the company’s business operations and existing data architecture.

Data Ingestion: Current Pipeline vs. With Airbyte
⚠️ Current Pipeline
Manual & Fragmented
💻 Woo
→
⚙️ Custom Script
→
🗄️ Data Warehouse
📦 Warehouse Partner
→
⚙️ Custom Script
→
🗄️ Data Warehouse
☁️ Other Sources
→
⚙️ Custom Script
→
🗄️ Data Warehouse
VS
✅ With Airbyte
Standardised & Scalable
💻 Woo
→
📦 Warehouse Partner
→
☁️ Other Sources
→
🔄
Airbyte
→
🗄️
Data Warehouse
Figure 2 — Airbyte replaces multiple one-off scripts with a single, standardised ingestion layer.

With the help of the company’s in-house chatbot and internal documentation, I was gradually able to piece together how the current data pipeline operates and why the company is exploring a more standardised ingestion solution. Through this process, I also gained a better appreciation of how data engineering decisions are closely tied to operational scalability and maintainability.

💡 Key Insight: Understanding why a technical decision is being made is just as important as understanding what the technology does. Context is everything.

Problem Solving

This experience reinforced an important lesson about problem-solving, especially when dealing with unfamiliar systems or technologies. Asking questions is extremely important. Instead of assuming understanding, it is often more effective to break the problem down systematically. Personally, I like to start with the “5Ws and 1H” approach — Who, What, When, Where, Why, and How. This framework helps me structure my thinking, identify knowledge gaps, and build a clearer understanding before attempting to propose solutions or write technical documentation.

My Problem-Solving Framework: 5Ws & 1H
👥
WHO
Who is affected?
Which teams or stakeholders are involved or impacted?
🎯
WHAT
What is the problem?
Define the task clearly — what needs to be done?
📅
WHEN
When is it needed?
Understand timelines, deadlines, and urgency.
📍
WHERE
Where does it happen?
In which system, service, or part of the pipeline?
🤔
WHY
Why does it matter?
Understand the business motivation and impact.
🔧
HOW
How do we solve it?
Propose a solution only after answering the above.
Figure 3 — The 5Ws & 1H framework: structure your thinking before jumping to solutions.