"Obviously this is just a sketch"

Remembering the origin of Omni’s semantic layer

steven-blog-hero

“That’s just a list of shit though!”

It was April 2022. The ten of us, Omni’s earliest employees, were ensconced in a conference room in the First Round Capital offices in SF for our first offsite. The eloquent speaker was me, addressing our CEO, Colin. I was coming off yet another terrible night’s sleep, after the birth of my second daughter. She was almost three months old, Omni was almost two months old, and as a (at that time) fully remote company, this was our first time all getting together.

And the list of shit? We were spitballing how to balance the governance of a semantic model (with which most of us had some experience) with a more flexible approach. Colin had just proposed that when users created custom fields while exploring, the product would remember those fields. These previously created fields would show up for other users in the field picker. Ten definitions of “something something revenue”: pick your favorite. That was the list. In that moment, I could not summon the words for a better critique.

This is the (very short) story of how we got from the list of shit to the start of our governed but extensible semantic layer.

For me, it started before that offsite, with Model.kt. After taking the first month of Omni’s life off, I had to get things going. Having been deep in the core of the modeling layer at Looker, I was the natural “semantics person”. I sat down and spiked on my real first commit. It was Github PR #22, “Initial semantics work”. It included “Model”, “ParsedSqlExpression”, “QueryRequest”, “AlgebraFromApi” and “GenFromAlgebra”. 

At its core, the semantic model was a data structure, not text. SQL snippets in the model lived as structured “ParsedSqlExpression”s. A query request with some fields in a base view was passed off to the “AlgebraFromApi” class, which used Calcite, the open-source SQL framework, to compose the ParsedSqlExpressions into relational algebra based on rules. “GenFromAlgebra” boiled that abstract relational algebra into actual SQL that could be executed.

The doc comment in that initial Model.kt read:

/** Obviously this is just a sketch. The big design decision inherent here is to go all-in on Calcite at this level,

 *  rather than trying to have an abstraction between model and Calcite. In my past life doing this, the abstraction was just annoying

 *  and forced all this black-boxiness, so if we're sold on Calcite I think it's worth doing this deep integration.

 *  Otherwise, you need to re-parse the SQL every time, or create some extra thing to hold SQL parses. */

To a remarkable extent, that basic proof of concept of Model/ParsedSqlExpression/QueryRequest still powers how the core of Omni query generation works today. 

Those fundamental classes are still there, still doing the same jobs. We’ve just added an incredible load of functionality onto them. Ultimately, nailing a good Calcite-based semantic query generation architecture wasn’t that shocking, as I had just spent three years rewriting Looker’s query generation layer to use a Calcite-based, Kotlin data-structure-first approach. But this nice little design skipped over one of the hardest parts of the “Model” class. 

Sure, with this framework we could build a great engine to take a semantic model and generate SQL queries. But where does the semantic model come from?

A few weeks later, we found ourselves in the list-of-shit conversation, discussing how to actually generate that semantic model. We thought Looker had nailed governed BI, but fell short in allowing flexible analytical querying. You could define custom fields in lexp, but they were restricted to scalar expressions of other fields. Merge queries were another bolt-on: you could join certain results together in postprocessing. “Table calculations”, also written in lexp, ran in a postprocessing step outside of the database as well.

Other tools had built richer analytical toolsets, but were missing governance and in-database capabilities. We needed to figure out how we were going to fulfill this vision to provide the best of both worlds.

I already had an idea. We would use a blend of a SQL runner and field picker and parse the raw SQL into new fields and model structures that would attach to the query in a “query model”. This plan turned out to be a bad idea for that moment, though it would have its day later; that’s a different story.

However, to get that parsing working, we needed some way to combine that query model (the little extension we parsed out of the SQL) with the main model, to actually generate the query. In PR #71 “Get basic parsing working”, I wrote a “Model.extendFrom” function. In the spirit of the “sketch” comment in the same class, it contained the following TODO:

// TODO: this is not really how we should do extensions!! Not performant to re-do every time necessarily and generally

// should be more thought out. But need to make some initial stuff work.

At some point during the “list of shit” discussion, Nate, another one of our founding engineers, spoke up. He had been looking at how “extendFrom” worked, and had actually built a bit of functionality on top of it in the initial “expedition” UI (an early version of “workbooks”). 

“What if”, he asked, “we used extendFrom for everything?” 

What if users could define their fields in a model they could easily edit, and have it extend over the main model? (The words “workbook” and “shared” came later.) What if the database schema could live in its own layer, which could stay current with the database, without interfering with the curated, governed model?

I don’t remember what I thought of this exactly or how others reacted, but I remember that moment, because that’s what we ended up doing. 

In that moment, Nate had just invented Omni.

Memory is funny. My recollections from that first offsite stand out clearly. Then, the subsequent weeks and months of work, where the actual product emerged, all fade into each other. Most of what I can reconstruct comes from “git log” and Google Calendar.

steven-blog-timeline-vertical

Everyone was remote initially, and worked from their home offices. At that first offsite, as an experiment, the company bought everyone Facebook Portals, with the idea that we would keep the device on all day and interrupt each other when needed, to compensate for the remoteness. This did not turn out to be a good idea.

Instead, everyone on the team went very heads down on their own areas. There was a meeting at 11am every day: team-wide and hour long on M/W/F, with a quick engineering standup on T/Th. Presumably folks partly worked during those full hour meetings, but again, the blur, I don’t really remember.

In those meetings, the paradigm emerged, gradually. Luke, Jonny, and Chris built out that vision of generating the schema model from the database. First, it happened on app boot, then through an API and a button.

Remember: there was no textual representation of the model outside of its raw JSON. If modeling needed to happen, it had to happen in the expedition. Richard built the custom field editor and the join modal. Luke fleshed out the initial quickagg vision (right-click a field, pick an aggregation, and get a new measure) and built way more. 

Initially, the model produced in the expedition was the only extension to the database schema model. If you started a new expedition, you got a new model. At some point, we decided to rename “expedition” to “workbook”. Alisa built the filter UI. Jared made it all look good.

Around July, we started trying to make the modeling layer more real. I built out a converter that turned the data structure into YAML and could save YAML back to the data structure, including in the “combined” mode, which shows a workbook’s changes merged with the shared model beneath it. 

Then, in one of the more surprising demos in Omni history, Jamie, co-founder and company president, dusted off his old coding skills to invent the whole workbook-to-shared promotion flow, completing the paradigm we imagined at the first offsite. Workbook models would inherit the shared model (the curated, governed one), and the shared model would inherit the schema model. Modeling could happen in the workbook, and be promoted to shared. Modelers could write and view their model in the YAML IDE.

steven-blog-layer-diagram

That was the answer to the list of shit. Fields you made while exploring stayed in your workbook, and only reached everyone else by being promoted into the shared model.

“extendFrom” lost its TODO: it turned out that merging these objects in Kotlin was tractable, and we built a memoization layer for large models. The “Obviously this is just a sketch” comment is still there in “Model.kt” as a sort of monument to the founding of the semantic layer. 

All the pieces were finally in place. Sort of.

There were no Topics. My initial code (which still kicks in when you use “All Views and Fields” mode) calculated optimal join paths from the graph based on direction and depth. In late 2022, we threw in the towel and realized users needed predefined bundles of views and the ways they join together. Sarah, who had joined a few months after that first offsite, came up with the term “Topics”, and we built those out.

There was no version control whatsoever. That didn’t come in until Cathy, who joined as my semantics partner in crime mid-2022, designed and built out an elegant system for recording up and down JSON patches (small records of what changed) on model save, applying them forward or backward to rewind and replay history of the model.

And there certainly was no git integration. With the semantic model being structured, not text-based (the YAML was produced and ingested ephemerally, to edit the model, but never persisted), git was a struggle. A year after that first offsite, when the whole team got together in person again in April 2023, there was a lot of hand-wringing about git and whether we could ever really make that work. It would not be until 2024 that Conrad, who joined in late 2023, pushed that over the line.

It’s hard to write about those early years without some nostalgia. This might strike non-software engineers as odd. It wasn’t that long ago! But thanks to agentic coding, the way we do our jobs has completely changed. All those days that ran together in that first year were mostly spent struggling through a way of writing code we simply don’t do anymore.

And yet, that’s not really the moral of this story. This story was about ten humans who sat together around a table in 2022, trying to figure out what they would build, and how. That remains the central question for anyone doing anything today. (Go ask a frontier model, “How should I build my startup?”)

It so happens that all ten of those people in that room, the initial ten Omni employees, still work at Omni today. I believe this is not the norm, and reflects something special about the company. Partly it’s due to hiring that initial team of known “good” (both in the sense of being productive employees and just genuinely decent) people. Partly it’s because the company has been successful. But I think one big part of it is something revealed by this story, something that Omni had at the beginning and has retained. It’s a place where we try out ideas: “expeditions”, “All Views and Fields”, day-long Facebook Portal calls, the “list of shit”. It has allowed us to adapt and feel empowered to keep making big changes, in response to a tech world that is changing faster than ever.

It’s a world that won’t slow down any time soon. There will be more decisions to make, more blank pages to fill in. You have to start somewhere, and so you might as well make a sketch. You never know how far you might be able to run with it.

We love tackling problems like this at Omni. If you're interested in building with us, we'd love to speak with you.