Writing · July 26, 2026

I built a lake


Three months turning HR data scattered across systems into an HR Data Lake anyone can draw from — within their own permissions.

#HR #AI #Data #Decisions

Hover to magnify, click to open full size

I built a lake,
and its surface is always clear.

About two months ago,
I arrived at an unfamiliar settlement
and took on a major task: to find a water source.

The settlement was made up
of many different plots of land.
Each plot had its own
elevation, geology and history.

Some plots held abundant water,
some were easy to dig,
some even bore the marks of earlier wells.

But no single plot
could meet every condition at once —
a stable supply of water that was safe, clean and worth trusting.

I picked up a pickaxe named Gemini
and started digging a SQL Server plot called “the recruitment database.”

Then the CHRO drove in an excavator named Claude Code (model: Max 20).

He tossed my pickaxe aside
and said: “Dig. The company is giving you the resources.”

So I began to dig.

A month later,
water welled up from the strata on its own.
Only, the quality wasn’t what I’d hoped.

I tried DBeaver and MSSQL,
taking the soil and rock layers apart one at a time;
I also started talking to the locals,
to understand what had happened to this land before.

I cleaned the data once, twice, three times…
past twelve times in the end.
Still no truly clean water.

In the second month,
the excavator turned to the next village.
Not far away sat another SQL Server plot: “the performance database.”

I tried, from different angles,
to draw water that was clean, stable,
and still held its informational value.

Then the CHRO said, just as casually:
“Why not build an MCP Server to filter the water?”

Suddenly it was clear.
The question was never only:
“Is the water clean?”

It was:

Do we have a way
for every person to draw exactly the water they need, within their own permissions?

So the third month began.
Between the two plots,
I built a collection station.

Through role-based access in Google OAuth,
five colleagues in five different roles
each draw a different level of data,
according to their own permissions.

The MCP Server’s filter bridge
means every act of “drawing water”
has rules, has boundaries, and leaves a record.

Finally, Claude Code CLI serves as the tap. Colleagues don’t need to learn SQL, or know which table the data hides in.
They just need to say, in plain human language,
the question they actually want answered.

On the other side of the lake,
every morning at 09:00,
the Windows Task Scheduler starts automatically.
It extracts the latest, de-identified data
into the lake in Apache Parquet format.

Then, through DuckDB queries combined with LLM analysis, the data flows out of the lake into different use cases.

At the same time, it automatically produces:
・the CHRO’s daily situation room
・a data-lake health report
・a data-access audit report

Further out on the lake,
I’m starting to see the shape of the next stage.

From Data Lake,
flowing to Data Warehouse.

And from Data Warehouse,
settling into purpose-built Data Marts.

The same body of water,
through different filtering, sorting and channelling,
reaching different decision scenarios.

And that’s it.

The HR Data Lake —
a live HR data-query system —
had arrived.

Now you can ask it:

“Why did turnover rise this month?”

“Which business unit has the highest talent-attrition risk?”

“Pull together this year’s performance-review results for me.”

“Through the lens of Dave Ulrich’s HR Value Proposition,
what is our talent strategy missing right now?”

You can even ask:

“If I were the head of this business unit,
what talent decision should I be making right now?”

It won’t cross data boundaries,
and it won’t ignore role permissions.

It simply takes data that used to be scattered
across different systems, databases and reports,
and reorganizes it into an answer a person can understand —
and think further with.

This is roughly what I’ve come to understand as:

AI × HR × Data

AI isn’t only about writing an email for HR.
It isn’t only about making a deck for HR.

It’s about giving HR the ability to move
from “describing what happened,“
to “understanding why it happened,“
and finally: “what should we do next?”

While building the lake,
I remembered something the CHRO once told me:

“Data used once produces information;
information, once analyzed, produces knowledge;
and only a great deal of knowledge gathers into wisdom and a way of deciding.”

I built an HR data lake.

Not far off,
I can faintly see what it will become:
a talent-investment decision-support system.

Above the surface
is the data we can see.
Below the water
is organizational memory, built up over years.

And as the water runs through different channels,
settling and sorting along the way,
what it finally reaches
is every person who has to make a decision.

What matters
was never how big the lake is.

It’s whether, when we need to make an important decision,
we can draw from this lake
a single cup of truly clean water.

And its surface is always clear.
The air by the lake is always still.

Read the original on LinkedIn