2017 • Data ScientistRecommender systems at Trackuity.
2019 • ResearcherOpen/Linked data at Ghent University - imec.
2021 • Software EngineerGPS and map data at TomTom.
2022 • Data EngineerHotel data at Lighthouse.
The visuals are AI generated.GPT 6 Astra is weird.
Fun opinions ahead.They do not represent my employer.
I have sources though.Links to Open Access papers. Not always the first publication though.
To understand the process of discovery.
To understand the process of failure.
Telling historical stories is the best way to teach.
To learn how to cope with life.
Let’s Not Dumb Down the History of Computer Science
Communications of the ACM · 2021
01 Introduction
02 History
03 Google
04 Observations
05 Cases
Future users of large data banks must be protected from having to know how the data is organized in the machine (the internal representation).
This paper is concerned with the application of elementary relational theory to systems ... of formatted data.
Highly parallel database systems are beginning to displace traditional mainframe computers
Ten years ago the future of highly parallel database machines seemed gloomy
They describe:
Data warehousing and on-line analytical processing (OLAP) are essential elements of decision support
Typically, the data warehouse is maintained separately from the organization’s operational databases
OLAP operations include rollup and drill-down along one or more dimension hierarchies, slice_and_dice, and pivot
It is impossible to reliably provide atomic, consistent data when there are partitions in the network
It is feasible to achieve any two of the three properties: consistency, availability, and partition tolerance.
It provides fault tolerance while running on inexpensive commodity hardware
The largest cluster to date provides hundreds of terabytes of storage across thousands of disks
The Google File System
map function that processes a key/value pair to generate a set of intermediate key/value pairs
reduce function that merges all intermediate values associated with the same intermediate key
The runtime hid all the partitioning, scheduling, networking, ...
MapReduce: Simplified Data Processing on Large Clusters
GFS has a relaxed consistency model that supports our highly distributed applications well but remains relatively simple and efficient
As most of our files are append-only, a stale replica usually returns a premature end of chunk rather than outdated data.
MapReduce + GFS did not guarantee consistency nor availability
Google made it work for their use cases
Origins of Big Data and NoSQL (through BigTable)
Running aggregation queries over trillion-row tables in seconds.
A novel columnar storage representation for nested records.
Dremel: Interactive Analysis of Web-Scale Datasets
It is the first system to distribute data at global scale and support externally-consistent distributed transactions.
We should no longer depend on loosely synchronized clocks and weak time APIs in designing distributed algorithms.
If you read any paper -- make it this one
Spanner: Google's Globally-Distributed Database
Built at Google to support the AdWords business
Scalability of NoSQL systems like Bigtable, and the consistency and usability of traditional SQL databases
F1: A Distributed SQL Database That Scales
The same tools are used for all data volumes
The only limiting factor is your wallet
Keeping costs down is arguably harder than before
Google needs your money to fund their AI
The Hadoop/Spark crowd met with the database crowd, both crowds often build data warehouses
We went from NoSQL to SQL abuse
Even Databricks isn't using Spark for their raw SQL products anymore (replaced by Photon)
Spark still excells for non-SQL workloads though, but just use Scala then
It is tempting, if the only tool you have is a hammer, to treat everything as if it were a nail
The Golden Hammer thrives in organizations where software teams fail to invest in education
AntiPatterns: Refactoring Software, Architectures, and Projects in Crisis (Brown et al., 1998)
Just because you can express a problem in SQL doesn't mean you should
SQL Antipatterns: Avoiding the Pitfalls of Database Programming (Bill Karwin, 2010)
Beware of the Turing tar-pit in which everything is possible but nothing of interest is easy.
A language can be universal yet offer little help with the task at hand.
Just buzzwords
Both terms refer to data warehouses: do you let the data warehouse handle the transformations?
Platforms like Databricks blend the boundaries so much it's a useless distinction
Never made sense outside of data warehouses, and therefore, outside of data analytics
Postgres+DuckDB query engine
pg_duckdb brings analytical queries inside Postgres.
SQL catalog→Parquet lake
Multi-table ACID transactions; small writes can stay in the catalog.
UPDATE→patch parts→SELECT
Read updated values before background merges.
Lakebase→open storage←analytics
LTAP: one logical dataset, specialised Postgres and analytical engines.
OLAP and OLTP on a single copy of data in the lake, eliminating ETL, replicas, and pipelines by design
Databricks is the world's first LTAP platform
No more copying of operational data to data warehouses
Google did it first with Spanner Data Boost
Processing the largest database in the industry (over 140 terabytes daily) allows Lighthouse to deliver global insights on pricing
Mostly HTML, and rest is mostly XML
'Modern' data tools bill per byte
Majority of transformations are Python running on Kubernetes
What is the price?
AI is great at extraction but too expensive
AI is also great at generating regular expressions
Streaming parser using xml.etree
xml.etree
Transformations in DuckDB and Polars (why both?)
Parquet writtem to GCS, merged into BigQuery as external tables
OpenStreetMap (256 GB RAM)
Proprietary map data
Massive amounts of GPS data
Sensor derived data
Not convinced these issues are fixable
Value per trace is low; costs per trace must be low
Item data from marketplaces(Immoweb, VDAB)
New items every day andnewest items are most valuable
Vague addresses
Edited images(banners, cropping)
Nothing like cryptographic hashes
Locality Sensitive Hashing (LSH)
Perceptual hashes can be as simple as thresholds of average luminosity
Companion to the Golden Hammer: overusing a favourite tool is a different problem from having expressive power without useful abstractions. Tie this back to tool choice; this is not a blanket claim that SQL is a Turing tarpit.