Episode 18

Johan van de Werken (Superform Labs, GA4Dataform, Select Star, Simmer) on Why Trust Starts at the Data Layer, Knowledge Compiled as Code, and the Hidden Cost of AI-Written BigQuery Queries

October 5, 2026 · Knowledge Distillation Podcast
Listen on Spotify Apple Podcasts Download MP3

Key Takeaways

  • Trustworthy AI analysis starts with a clean, shared data model
  • AI lets people do more - and break more - so QA matters more than ever
  • Unbounded AI-written BigQuery queries can cost cents or thousands
  • Context and lineage files let agents trace where every number comes from
  • The durable skill is knowing what your metrics actually mean

In this episode of Knowledge Distillation, Katrin Ribant goes underneath the analysis and into the data layer with Johan van de Werken - co-founder of Superform Labs, the team behind GA4Dataform, the open-source framework that turns raw GA4 BigQuery exports into clean, sessionized, ready-to-use tables. GA4Dataform is also built into Ask-Y’s GA4 BigQuery connector, so this is a conversation with someone whose work sits right inside the product. Johan explains what the framework actually does, why a shared starting point for GA4 data makes sense when every dataset has the same structure, and why the real work - your business logic - only starts once the boring part is done.

From there the conversation turns to what AI changes for people who work with this data. Video courses are giving way to live teaching, and Johan’s team now compiles its knowledge as code, with context files that teach AI agents how to read the models, how the tables relate and what your company’s own logic is. Then the warning every analyst using an LLM on BigQuery needs to hear: a single day of GA4 data can be billions of rows, so an unbounded AI-written query - or a dashboard sitting on a granular table - can cost anything from a few cents to thousands, or worse, return a confident but wrong answer.

They also cover how roles are shifting: clients can now do far more themselves, which means they break more, which makes QA and an agreed ground truth more important than ever. Johan shares why GA4Dataform classified organic AI traffic long before GA4 did, what that classification can and can’t capture, and what he sees as durable for analysts: understanding what your metrics actually mean - what is a session, what is a user - well enough to judge whether an answer makes sense. And maybe, as he puts it, he’s still an educator - only now he’s teaching agents.

All episodes on our website: www.ask-y.ai/knowledge-distillation-podcast

Learn more about ASK-Y: www.ask-y.ai

Chapters

  1. 00:00 Introduction
  2. 01:53 Meet Johan van de Werken and GA4Dataform
  3. 03:44 What GA4Dataform Is and the Problem It Solves
  4. 06:49 A Shared Starting Point for GA4 BigQuery Data
  5. 09:53 Community, Core and Premium: Attribution Models and User Journeys
  6. 13:30 Education in the Time of AI
  7. 16:38 Knowledge Compiled as Code: Context Files for AI Agents
  8. 19:19 BigQuery Costs and AI-Written Queries
  9. 26:29 From Data Analyst to Analytics Engineer: How Client Roles Are Changing
  10. 30:26 QA, Ground Truth and an Open-Source QA Framework
  11. 35:11 Why Dashboards Aren't Trusted - and Traceability With LLMs
  12. 38:57 Lineage Files and Trusting the Source
  13. 41:02 Classifying Organic AI Traffic
  14. 43:34 Paid AI Traffic and Agentic Commerce
  15. 46:22 Will AI Take the Analyst's Job?
  16. 50:02 Durable Skills: What Is a Session? What Is a User?
  17. 53:10 Teaching LLMs vs Teaching Humans
  18. 57:49 Where to Find Johan + Closing Thoughts

Resources & Links

Companies & Tools
  • Superform Labs – the team behind GA4Dataform, co-founded by Johan van de Werken
  • GA4Dataform – open-source framework that models raw GA4 BigQuery data into sessionized output tables (Community, Core and Premium versions)
  • Select Star – Johan’s analytics freelance practice in the Netherlands
  • Simmer – Johan’s GA4 BigQuery video course
Analytics & Data
  • Google Analytics 4 (GA4) – raw event export to BigQuery
  • Google BigQuery – discussed in the context of data volumes and query costs
  • Dataform – Google Cloud’s data transformation workflow tool that GA4Dataform runs on
  • Looker Studio (Data Studio) – referenced for dashboard query costs on granular tables
AI
  • Claude, Gemini, ChatGPT, Perplexity – referenced as LLMs and as sources of AI-referred traffic
  • Ask-Y Prism – AI analytics platform with GA4Dataform built into its GA4 BigQuery connector
  • Ask-Y open-source QA framework – ground-truth QA for analytics changes, available on the Ask-Y website
Previous Episodes Referenced
  • Simo Ahava – Episode 14, on education in the time of AI
  • Kelly Wortham – Episode 17, on AI-referred traffic and agentic commerce

About the Guests

Johan van de Werken
Host
Katrin Ribant

Host: Katrin Ribant Guest: Johan van de Werken Podcast: Knowledge Distillation
Episode: 18 Runtime: ~59 minutes Release Date: 10/05/2026 Editing: Tom Fuller

Full Transcript

Katrin (02:02)
Welcome to Knowledge Distillation, where we explore the rise of the AI analyst. I'm your host, Katrin Ribant, CEO and founder of Ask-Y. This is episode 18, and today we're going underneath the analysis into the data layer. the layer that decides whether anything is trustworthy in the first place. And whether we can actually look at our data and get answers. So I have a slightly unusual reason for having today's guest on. It's an absolute pleasure to have you. Johan van de Werken is the co-founder of Superform Labs. Welcome, Johan.

Johan van de Werken (02:02)
Thanks.

Katrin (02:41)
Superform Labs is the team behind GA4Dataform, an open source framework whose flagship product is a set of methods. To aggregate GA4 BigQuery data. Anyone who's ever worked with GA4 BigQuery data knows that getting anything out of it means the first step is you have to create a data model. That's what GA4Dataform form does for you. And at Ask-Y, we've built GA4Dataform into our GA4 BigQuery Connector. It's the foundation of our AI analyst read of how AI analysts read GA4 through. So This is a conversation with someone whose work is sitting right inside our product. first of all, thank you for doing that, Johan. you really you really did something very big for the industry there. It's a pleasure to have you on the show. You do so many things beside working on GA4Dataform, but before anything else, for the people listening who have never opened Dataform in their life or never actually work with BigQuery data, GA4 BigQuery data. What is GA4Dataform? What problem does it solve? And why did you guys decide it that it should exist?

Johan van de Werken (03:56)
Excellent question. so yeah, as you said, every GA4, so Google Analytics four user can export the raw events they collect on their websites or in their apps for free to BigQuery, which is Google's data warehouse. By default you will get, depending on your traffic, a small or an im enormous amount of rows of events and their data is nested, it's complicated, there are errors in their definitions you probably are not aware of when you are just using the user interface of GA4. There are no sessions to be found, only events. So to report to use the data to, you know, create a create a report or well talk with your AI agent. to your GA4 data. You first need to sessionize the data or aggregate it on user level or whatever your use case is. And that requires a lot of knowledge. That requires that you're able to write SQL code. and you need ways to do that to make it update every day. So every new day of data coming in also refle is reflected in your output tables that you you use for your reports. So to do that we we use Google's built-in tool in Google Cloud called Dataform, which is basically service to manage your data transformation workflow. so when the data is collected you can use code to take that data and transform it in any way you want. Let it produce output tables and and do that every day or every hour to have an updated pipeline of your GA4 data. And well GA4Dataform is is actually that code inside that tool that you can install. Because we also built an installer around it, you just log in, click a few buttons, we install the code in your own Google cloud environment, and from them it takes the raw data you already have in your BigQuery environment and based on a schedule you set it transform transforms that data into ready to use output tables, basically. Does that make sense?

Katrin (06:49)
It makes total sense. And so in doing this, you essentially set an industry standard for how to sessionize GA4 BigQuery data, right?

Johan van de Werken (06:49)
Yeah. Yeah.

Katrin (07:03)
and I would say that's probably a de facto standard, which is as much as we get in digital analytics anyway, but I would say it's probably at this point the universal standard for for do we doing this work. And You know, obviously everybody can use it because it's open source, it's public. what do how do you think about it when you think, okay, so we set the standard and we and you're basically maintaining the the the standard. what does it r require from you and what does it stop you from doing when you update the package?

Johan van de Werken (07:42)
So standard is obviously a big word. We we we promote it like that. and it's also a bit true because you we we are with five basically five freelancers exactly doing this stuff every day. And every day we we we build those models ourselves. Every time slightly different. So it made sense, at least for us, for our own consumption. To create a standard and and of course because the data set is always the structure of the data set is always the same for everyone. So it made sense to create something that everyone can use. And and therefore we we call it a an industry standard. That's of course up to the users to to define if it's really a standard. But I what I really mean is a a a s a shared starting point because What we provide is only the boring stuff. It's only the stuff everyone needs to repeat to make sense of the raw events data, to create a bit of context, to add context, session context or user context. And the real fun starts when our when the when the results of our tool is done, then the real fun starts for the user because then you can add your specific, you know, organization or business logic to it. And that's something we cannot standardize for you.

Katrin (09:12)
Yes, and obviously that's something where a platform like Prism come comes in place to help with that. but you know just to to hone a little bit more about around what you do with with GA4Dataform, one of the things I want to talk about a little later is how you did a classification for AI traffic and what you're thinking was behind that. But also there's an aspect of this as this being a standard where so you have this core offering and then you're building around it. Can you tell us a little bit about how you're building around it, how you have a paid offering around it, and what it is that you're adding and continuously adding to the packages.

Johan van de Werken (09:55)
So indeed we have a to be precise we have three versions. We have the open source, which we call community, that's really a GitHub repository. You can read the code, you have to install it yourself. Then we have the core version, which is that open source code, but you can use the installer we built so you don't have to deploy the code yourself. And then we have the premium offering and that includes a lot more use features. we call it modules because it's a modular approach. So GA4, the standard sessionization and the the unpacking of the raw data is it's like one module, but we also add other modules like and and usually we do it per use case. So recently we added the attribution module. So that's that's something that's off by default and you can enable it and what it does is provide a lot of extra tables again in the same workflow so every day those tables will update as well all using the raw data and for instance the attribution module provides six attribution models first click last click time decay linear all those good old attribution models that used to be in Google Analytics and User Interface. But nowadays with GA4 I think there's only two or three left because they they remove the other ones. So and the fun part around is that you can configure them, you can change their, you know, how they are built as well in the configuration files that we also include. So And based on that you can also create, for instance, a a table that shows you the journeys of a user. So let's say they have different touch points, as long as there is a user ID, and they have different sessions coming from different marketing channels, we can show the the actual journeys. That's all that data is already collected, but it's not It's already in there, but it needs to be you know it needs to be the logic that is needed to harvest that data, that's a lot of work. It's very specific knowledge

Katrin (09:55)
Yes.

Johan van de Werken (12:35)
and we can that's the part we can standardize.

Katrin (12:39)
Exactly. So like that's kind of like where I wanted to get to is you guys have a really staggering amount of knowledge about this particular process, this niche, and the industry in general, from the data layer perspective. And you bring a really, really fantastic amount of care to the way you are building these these methods. putting some of it, you know, at the disposition of the community, allowing tools like ours to you know, to not have to redo all of that. So first of all, thank you. but also you will obviously all of that knowledge is out there because you publish it. And you also have, and I know I'm sort of walking into into because we we had this conversation right before the podcast about about how you consider yourself being an educator or not. I always thought you consider yourself being an educator because you do obviously the the the BigQuery course for Simmer. So I had Simo as a guest on the show about six months ago, and one of our topics was education in the time of AI, and what his point of view was about the need to learn and how to teach in this new world where so much of the information is available in any form, any in any sort of modulable for form through AI. What is your point of view? I'm just gonna say, as an educator, in the time of AI?

Johan van de Werken (14:17)
Yeah, it's it's it's I I think I I share the experience and also I think Simo mentioned it in one of his recent newsletters that the demand for online video courses specif specifically around these topics is declining as well. No surprise to me. I I did expect that, Simo as well, I guess. because one upside of you know, the the way an LLM can help you is to help you more better with your specific use case because to transfer knowledge in a video course you have to assume you have to do assumptions about your audience and because you you don't know everyone, you need to keep it as generic as possible, at the same time as detailed as possible, but there's a balance there and it's it's a hard one. You cannot go into everyone's specific use cases. That's impossible. So therefore I think the move they are making with Simmer is a is a is a smart one to to lean more into live courses where you can specifically help students with their questions. So with a select group, maybe ten people. And then it's possible to listen to everyone and to to teach them stuff about their own use case. Of course that's not possible in an on demand course. So that format for me is basically dead.

Katrin (16:00)
And so the format is dead, but not the need for knowledge. in in the, you know, in in the obviously knowledge from, you know, specialists like you, but also knowledge for practitioners, right? Because yes, you can get a lot of help from AI, but you really do still need to know what you're doing, and know what the AI are is doing and one of the implications of knowing what you're doing as a human when you handle the type of data. we work with specifically when it it comes to like, you BigQuery and the larger amounts of data. Any ba anybody

Johan van de Werken (16:00)
So

Katrin (16:37)
who yes, sorry, yes, go ahead.

Johan van de Werken (16:39)
the the way I see it is that GA4Dataform is is that knowledge, but then compiled as code. And also we added instruction files for agents to to learn the agents how to read it, how to treat the data, how to treat the code. That's that's that's the way I I approach it now. So all the knowledge we still have we want to transfer it, but then we can use our own tool to do so, for instance. And and probably you can do it with your own tool as well.

Katrin (17:18)
And so so you have that Gemini.md file, right? That's what that's what you're referring to here.

Johan van de Werken (17:24)
Yeah. And also for the other LLMs, yeah.

Katrin (17:26)
So it's one file for all the LLMs or you have several files per LLM? Are you treating them differently? Are you doing something something specific for any of them?

Johan van de Werken (17:37)
I think we do. So Chris, one of my colleagues, is well, he's the AI guy. He's basically in turning into an agent himself nowadays.

Katrin (17:37)
Yeah.

Johan van de Werken (17:50)
he knows all this stuff. I think mostly it's duplicated, but because some LLMs require their own files, right? Then we just added added multiple files. But but the most important thing is we also included a file, I think it's called my my company or something, dot Md where you can add your specific business logic or business knowledge to GA4Dataform as well. So it it does not only know about GA4Dataform and all the code it contains and the logic, but also it knows something about about your own company. So that's for the you know the the clients or the users that only use GA4Dataform to produce the output tables and not some other data transformation layer for their own models. So it's all about providing the context I would say.

Katrin (18:59)
Yes, it definitively is with like with AI it is it is really all about providing the context and making sure that the the LLM can pick it up and different obviously different LLMs pick it up differently. And so on the on that subject of cost, at some point, I think it was in August, you wrote this whole walkthrough as response to a question you had about whether your incremental pattern forced a full table scan on every run and potential potentially then would like run up the cost. of querying

Johan van de Werken (18:59)
Mm-hmm.

Katrin (19:27)
as they go by. This go by. And and you basically exp published a whole whole walkthrough explaining how it works and optimizing and like basically pointed out that the question was assuming a different cost model in BigQuery than what was the real cost model. If I remember well, that was kind of the crux of of of the confusion

Johan van de Werken (19:27)
Mm-hmm.

Katrin (19:53)
there. Obviously people do work in BigQuery People also transfer that data into different data warehouses. All all the different data warehouses have different cost models. At this you know, at this at this level of of query querying, we we're talking about something that can have really potentially large impacts. When you think about knowledge About just general knowledge about working with this complexity of data, which you're obviously specialist in. What do you think that people today should think about when using AI to generate these types of queries? Do you guys internally like have developed some sort of methodology, some sort of checklist that you that you're using when you are developing your types of queries, the types of queries that you put on the market?

Johan van de Werken (20:56)
Hmm. Internally I'm not sure. We what we what we typically do is work quite independently on our own features. So there's a lot of freedom for within our team to work on a feature. So we we don't before we start building, we don't have large meetings about them. But in general I would say most people are probably not aware of the sizes of the data sets that we're talking about here. That's very specific for the digital analytics industry, I guess, that tables can be really, really big when you have a lot of traffic, right? Yeah.

Katrin (21:35)
So just to give people a a a notion, you know, people who have never manipulated this type of data, when we talk about a lot of data and running up cost and the risk of running

Johan van de Werken (21:35)
Right.

Katrin (21:44)
up cost, what are we talking about in orders of magnitude in your experience?

Johan van de Werken (21:49)
So for instance, one day of GA4 data for an a b a bigger company can contain billions of rows. It can so one one table in BigQuery can be like twenty gigabytes. There's there's a lot that are way bigger than that. Usually for small websites it's it's a couple of hundred of megabytes. But you can imagine if you're If you just ask a very generic question to an LLM, you don't you don't specify a r a date range or you do and you just say give me that and that number for the last three years. And every the table for every day is I don't know, twenty gigabytes. You can imagine the query cost that will

Katrin (22:44)
I can, but I don't think most users can. I'm actually getting white just thinking about that. What what kind of numbers, what kind of cost are we talking about?

Johan van de Werken (22:55)
So BigQuery is actually pretty cheap when it comes to querying, right? Then but but you can so for I'm a freelancer as well. For my clients I am always cautious about this. I only quer query aggregated tables. So tables that are already that are the output of the data transformation layer, right? Don't I don't query the raw data set because it's it's too It's not efficient, it's it's too expensive. But it can I don't yeah, it can be it c it really can be anything from a a few cents to hundreds of Euros per query. It it really depends. So so you need some knowledge about it. Yeah.

Katrin (23:40)
Yes, I mean it ca it can get it can get it can get to thousands of dollars on a on a on a bad query. Like that's n definitively

Johan van de Werken (23:40)
Yeah. Yeah.

Katrin (23:45)
not unheard of. and and and if you if

Johan van de Werken (23:58)
And what people don't don't don't know, a lot of people don't realize that if you connect a data studio report to a BigQuery source table. and one hundred users are using that report that and they all change the date range, click some filters, every interaction with a dashboard is a query, right? So if the

Katrin (23:58)
Yes.

Johan van de Werken (24:23)
table underneath is not aggregated for that specific use case, let's say to aggregate the data by month. But underneath there is a very granular granular table on session ID level, for instance, which is not a best practice, but it happens a lot, then maybe one click can generate a query that needs to scan terabytes of data, let alone that hundred people click that button. So it's it's very important to be cautious about to think about those kind of things and and and all of this is not happening or very little in LLMs as far as I have seen. Again I'm not an expert on this, I'm not a data engineer. I use Claude as well. to query stuff. And I don't mind because I know the tables I use, the the tables I expose to Claude are already aggregated. So that will be okay. But you know if you if you don't if you are not cautious and you don't have a clue about what Claude or any other LLM is doing, then you can be surprised at the end of the month.

Katrin (25:52)
Yeah, it's it's if s if some if you do something complex and somewhere in that SQL there is something that you they haven't realized is a large scan, and you just run that, that can be a really bad surprise. So

Johan van de Werken (26:09)
But maybe worse maybe worse, next to the cost you can also end up with a total bullshit result. And and that can also

Katrin (26:15)
There is that.

Johan van de Werken (26:16)
be very costly.

Katrin (26:23)
There absolutely is that. so so you you obviously you freelance, right? So you say you're not a data engineer, what would you define yourself as when you freelance? What's what is typically the remit you cover?

Johan van de Werken (26:34)
I would say anything from from a data analyst to an analytics engineer. And so analytics engineer is basically a fancy term for someone that manages data transformation workflows. But I also

Katrin (26:34)
And still

Johan van de Werken (26:48)
create dashboards and I also do data analysis.

Katrin (26:53)
And so I'm I'm actually really curious about how you've been finding that roles evolve in in organizations. Because obviously with AI, all of our remits have become wider, right? It's it's it's really an an moment where you can touch some I mean I'm not particularly great at SQL, but I can execute on workflows that I would not have been able to to to do before by myself. Similarly, there's some domains I really don't know from a business perspective. I can at least gain some understanding of what are the key things people optimize in XYZ industry that I have maybe never worked on or whatnot. and that's obviously the case for everyone. I have two parts to this question. One is has it changed anything to what you can execute on and what you are doing in your in your freelance practice? And do you see with some of the clients that you're working with either a compression of roles onto less people or a broadening, you know, broadening of everybody's role? Like do you see any evolution in how people actually work from, you know, from the analyst to the marketer?

Johan van de Werken (28:13)
Yeah, so to start with my experience is that my clients I don't know if they become smarter, but they can do more technical stuff, right? So they they are actually changing let's say Dataform models I built, custom models. They're adding stuff to it and and publishing it and you know doing And that's and that's good because they they contain the business logic, right? They know which filter to add or which exception to to add. So so that's good, but they are not aware of technical complications that it can bring. So I would say my role is I I definitely need to add more monitoring, what to look what they are doing. It's of course it's their

Katrin (28:13)
yes.

Johan van de Werken (29:11)
project, but they can break stuff I built by for instance. So and probably if they need me, it's not about me having to add something because I teach them how to do that themselves with the help of, you know, AI. But I probably need to fix something because they broke something or they didn't they added something, for instance, or created a new model, but they are mixing up dimensions or metrics or and and that so they still need someone I would say to to look closely what they are doing to check if they are making no mistakes on a conceptual level.

Katrin (30:00)
That's really interesting that you you know that that you mentioned that because we so obviously w we also face the same you know the same sort of conundrum where people can do a lot more things, which means people break a lot more things. And so that requires a lot more QA. Nobody likes to do QA, obviously. And

Johan van de Werken (30:00)
Yeah. Nope.

Katrin (30:21)
and and and There aren't really very good methods to do QA because ultimately what you have to do is you have to sort of like have a trace of what is the ground truth, what is considered as being correct, which is purely a human definition. It's like this we considered as being correct. And then some

Johan van de Werken (30:41)
I would say common sense. Something like that.

Katrin (30:41)
Yes.

Johan van de Werken (30:41)
Yeah.

Katrin (30:44)
Well and and based on business knowledge, right? Because it and based on what is achievable with the data and what we will accept as errors, like it really is a human decision of to say, This is my ground truth. It's as good as it gets. This is what I need to get back to when I make a change. So I make a change and I know that I'm gonna change a certain number of things and a certain number of other things should remain the same. And now I'm starting QA. There isn't really a very good automated way to do this. And

Johan van de Werken (30:44)
No, no.

Katrin (31:17)
so because we have to do it a lot, we don't like to repeat

Johan van de Werken (31:17)
Yeah.

Katrin (31:26)
things too much. We've actually built a I would say an agent. We can call it more more or less maybe an agent. an open source sort of framework. It's also a GitHub repository that you can download and use in Claude or in any LLM that takes it. Takes the QA process, basically mimics the QA process from a very practical point of view. Because I find that everything that you have in tools like dbt or whatever, it has a level of technical QA, but it's not practical. It's not looking

Johan van de Werken (31:26)
Mm-hmm.

Katrin (31:56)
at ground truth and then com then there's a change and then comparing to ground truth. So basically the thing takes screenshots of what makes you establish what is ground truth. And then when you make a change, goes through everything and and and literally compares with with screenshots and with the you know, with like like splits in the data, what is correct, what is incorrect, what is changed, what hasn't changed. We hope that people are going to use this and contribute

Johan van de Werken (31:56)
Mm-hmm.

Katrin (32:23)
it to it because it's only gonna be good if it has practical experience. Like

Johan van de Werken (32:23)
Right. Yeah.

Katrin (32:29)
a lot of practical experience. So you know so that that's my shameless plug. you can go to our website and and and

Johan van de Werken (32:29)
Nice nice pitch.

Katrin (32:35)
and and yeah thank you and download that and and and it's it's completely free. It's literally just because it's so annoying to do QA that I feel like the world needs the world

Johan van de Werken (32:35)
It's it's boring, yeah. Yeah.

Katrin (32:49)
needs a solution for QA. So

Johan van de Werken (32:52)
Well that's that's exact that's exactly where why AI is great. It's to automate boring stuff, right? Stuff that

Katrin (32:52)
Yes.

Johan van de Werken (32:58)
no one wants to do. That's that's great. Yeah.

Katrin (33:02)
So so you know, so it's like r refreshing in a way, and validating to hear that you also face the fact that people can do a lot more things, which means they break a lot more things. And then a lot of your j

Johan van de Werken (33:13)
Yeah. I always use the the example of you know in in Google Analytics I'm not sure if it's still possible, but you were at some point, maybe even in universal analytics, in the previous version, it was you were able to to take a dimension and then add a metric to it. and but it it didn't match, right? Because

Katrin (33:13)
Yeah.

Johan van de Werken (33:37)
it's a different scope. Or something. Something

Katrin (33:37)
Yes.

Johan van de Werken (33:40)
is user scope, the other one is session scope or or event scope. And so

Katrin (33:43)
Yes. It's unr it's unrelated, yeah.

Johan van de Werken (33:45)
yeah, so the fact that it's possible to combine a dimension and a metric doesn't mean it makes sense to do it, right? Or to that the

Katrin (33:45)
Right.

Johan van de Werken (33:53)
result is is meaning something, or at least true or something. So that role for that and I think because I w what's the previous episode with with Simo as well on your podcast. I th I I think he's saying something similar, right? That you need someone at the end of the day with some I don't know what he's calling it, some spider sense, something like that. Like Yeah. Like

Katrin (34:24)
Yeah, I don't remember the world the word he means, but yes, I see exactly what you mean, yes.

Johan van de Werken (34:30)
like common sense just validating the results of someone else work. And and for that AI is is becoming better at that, but you know for for a while we are still I I think that's the direction of at least a part of my job. That that we need to validate other other work. Yeah.

Katrin (34:50)
And so then what do you see in terms of or do you see anything in terms of changes of remit, compression, expansion within the the organizations you work with, between data and marketing or within any of those functions?

Johan van de Werken (35:05)
So I don't know, in most in most organisations there's the confusion always is only getting bigger and bigger because dashboards are not trusted anymore. Usually that's what I see. And sometimes rightfully so because but but the problem is not getting fixed because with LLMs it's even more impossible to check the to check if the number the LLM is spitting out is true or where it originates from or how it is calculated, right? So So yes, it can be more specific about the question you are asking and and a dashboard, just like a video course, is more a generic attempt to answer a question. So yeah, the answer from your LLM feels more specific to your use case, but still how can you trust it? That's that's still an answer that's not solved. I think your tool is probably a possible solution to that, you know, to that yeah.

Katrin (36:17)
No, I I agree. This is this is definitively two like there there's really two main things that I think we we we are honing on you know like really strongly is one, traceability, like understanding this where does this come from? Because if you do like doing this purely in Claude, because you don't really have the I would say the substrate, the infrastructure behind it to organize the work, it's very difficult to trace anything.

Johan van de Werken (36:17)
Yeah.

Katrin (36:46)
To understand why some why why a response is the way it is, especially when you start doing long sessions where you have a lot of context,

Johan van de Werken (36:46)
Mm-hmm.

Katrin (37:00)
it's atten like the even with the really large models, the attention mechanism isn't good enough to pick up in the context, between the context and the memory, pick up the right information to give you a complete in a complete response. And so

Johan van de Werken (37:14)
Every every time my Claude is saying, Hey, this chat is becoming too long, I will I'm going to summarize it a bit to c to keep going. I'm thinking like, shit, now I'm going to lose a lot of details we need, right? So and that's ri literally what's happening.

Katrin (37:30)
And if you don't have a very, very good methodology to ask it when it gives you an answer, did you check this, this, this, this, this?

Johan van de Werken (37:30)
Yeah.

Katrin (37:39)
Because it it it just like at that point it just like forgets and and brushes aside the things that are too complicated or too too too token rich, you know, too too too token expensive. And then when you re ask it the question, it's like, yeah, I forgot that. no, you're right, I was wrong about this. So So if you don't if you're not very disciplined about about guiding it that way, which is, by the way, a lot of work. Like if you if you're doing real work with it, right? It's truly a

Johan van de Werken (37:39)
It is yeah.

Katrin (38:07)
lot of work. And it's actually also, I find, a lot of cognitive load. Because you have to keep trace of what is in every session, where every session is, where are the artifacts, where are the tables, where you are. Like it's it's kind actually really quite a big a bit of cognitive load. So one of the things we we you know we we're working on obviously is the memory and the context and serving the right context in the right queries. The other one is purely, like literally purely a workflow aspect of what what does you do you organizing your table, organizing your query, organizing your versions, organizing your artifacts, etc., into a tree, that means that you can trace things easily.

Johan van de Werken (38:07)
Right, yeah.

Katrin (38:46)
but I I agree this is something that is that is currently like If you actually work seriously with AI, it's very hard to do because you have to do all of that engineering yourself.

Johan van de Werken (38:57)
That's that's also a a context file we add in GA4Dataform. That's called lineage.md. That's for the agent to know how all the different models relate to each other, right? Why?

Katrin (38:57)
Yes.

Johan van de Werken (39:09)
So if if you need to change something on the in one of the output tables, it needs to know where that comes from and how it how it is aggregated from one table to another, to another, to another, table to another, right? So

Katrin (39:09)
Exactly.

Johan van de Werken (39:22)
so it it does it needs to quickly understand relations between between everything and so yeah so so I would say if if in doubt I ask in my process I ask the LLM to go back to the source always. So you need to be able to trust the source. And one of those sources is GA4Dataform and and the reason that that that source can be trusted is because all the humans put all their knowledge in it, right? And are

Katrin (39:22)
Right.

Johan van de Werken (39:52)
maintaining it. And and that's how I see it. So every tool needs their own GA4Dataform. Like if there's raw data coming from a tool, there needs to be like a a s a standard model, data model for it, right? And

Katrin (39:52)
Mm-hmm.

Johan van de Werken (40:17)
and obviously for the for the well known let's say you have an e commerce store and you run on Shopify, that's so well known that there are good models for that, right? But say you you build your own tool and you have some data sitting in there, you need to to to build it yourself because it's specific to built for you, right? So it's it's it's very it's a lot of work to be able to build to

Katrin (40:17)
Yes and and the the

Johan van de Werken (40:50)
build something with sources that you can all trust.

Katrin (40:53)
Exactly. Because in analytics, if you cannot actually trust all of it, then you can really trust none of it.

Johan van de Werken (40:53)
Exactly. Yeah.

Katrin (41:04)
And so let's let's go back to that organic AI classification that we talked about in the, you know, in the beginning of of the show. So I'm really curious about how you you decide to do this relatively early. I think it was 2025, right? So it was like very, very forward

Johan van de Werken (41:04)
Yeah.

Katrin (41:24)
thinking. first of all, I don't know if you see an evolution in the volume of that type of traffic. do people turn this on, do people use it? And what is your thinking behind what is getting classified as organic AI for you?

Johan van de Werken (41:43)
So people use it because I I think for new installations it's enabled by default. So only for backwards compatibility it's not enabled. we indeed introduced it quite early, way before GA4 itself introduced it. And for us it was just we saw traffic in, you know, in channels, referrals or unassigned. Clearly coming from ChatGPT mostly back then. I think maybe that was the only tool or perplexity, maybe adding UTM parameters, right? So then for us it make it made sense to create an optional extra channel group for it because we yeah, we want to stay close to GA4 definitions, but if we believe GA4 is missing something or doing it wrong, then we are then we just fix it for our users. So that's that's basically how we how we thought about it. it's of course it's dangerous but because the assumption when you introduce such a thing is that all maybe users think that all their AI traffic is now collected. Which is of course not the case because not all tools provide those parameters and so it's better than nothing, but it's it's not perfect, I would say. But what we did is just look at all the referrals and unassigned, look at the source mediums, and identify is this an LLM tool or not, and add it to a list. That's all we did.

Katrin (43:32)
Do you see anything happening with the I I I really have no idea how this actually works. So OpenAI introduced ads this year. and so now you technically have paid and I don't know how big this I I I think it's growing quite fast. at least in the US it is. I don't know about

Johan van de Werken (43:32)
Mm-hmm.

Katrin (43:54)
about Europe. do you do you see anything related to Bay traffic coming to the website from AI currently or not yet?

Johan van de Werken (44:03)
I haven't seen it yet. I also didn't look. I think this is more a a topic that's watched closely by people like Simo.

Katrin (44:03)
Mm-hmm.

Johan van de Werken (44:14)
So I don't know. Let's just say I don't know. Yeah.

Katrin (44:17)
I'll see him next time he comes on the show. the it's actually it's actually a topic that is is kind of like close to my heart because I would say probably the past six six episodes or seven episodes of the show have been around AI agentic traffic because I probably do think that as it grows and will grow, there is no way no it won't. As it grows, it will force people to reconsider the way that they build websites, definitely reconsider how they measure traffic, and also how they build KPIs around this traffic. And it seems, at least from the the episodes I did I did with Kelly Wortham a few months ago, it seems that AI traffic went from being low converting to being higher converting than non-AI traffic as the LLMs have become better at at pre-qualifying people and doing the research for them. So this is potentially extremely valuable traffic coming in that you want to treat completely differently than any other traffic because people come with a level of conviction and with the requirements about the website that are completely different.

Johan van de Werken (45:28)
Yep. But the question is, will they come in at all or do they check out in in OpenAI, in in ChatGPT, in the tool itself, right?

Katrin (45:41)
It it's a good question. they started with that and then they turned that off. So, you know, I

Johan van de Werken (45:41)
Okay.

Katrin (45:45)
don't I I I don't know what's going to happen with it. I imagine, you know, they're weighing car carefully whether they should turn it turn it back on or or or not. And and I I obviously nobody has any idea how that's going to evolve.

Johan van de Werken (45:58)
This is this is uncharted territory, right? Yeah.

Katrin (46:03)
completely. but it's fascinating. So

Johan van de Werken (46:03)
It is.

Katrin (46:05)
like, you know, like the speed at which this this evolves obviously is is staggering. And I feel like, you know, a few months ago there was a lot of fear in you know in the in any anybody who was working with anything technical, I would say, about, you know, AI taking your job, etcetera, etcetera. Where do you stand on this? obviously I you know I know I know what you said previously. I'm just going to let you repeat it if you want to or not. but where do you stand currently on AI taking your let's say your current job? Because obviously you come from from another job and you know

Johan van de Werken (46:51)
Yeah. Yeah, my my career is a mess. Thanks. so it's rich, yeah, a lot lot of experiences. Lot of experiences.

Katrin (46:51)
No, it's rich. My care my career is a mess too.

Johan van de Werken (47:03)
I so for in yeah, to start with I I wouldn't mind if AI took my job because it can be boring as hell as well. So let me just keep the fun part then. But Yeah. It's it's still too early to say I would say. I'm I'm not quite sure Will where we're heading. I think at at some point we will start to reevaluate and to reevalue the human aspect of things, the human the human eye or the human way of looking at things, people will eventually n so there's there's multiple points to it. From an economic point of view, it's it's not something that can go on forever because it all these tools cost a lot of money. And you know, at some point these companies need to ask much more money to for the usage we are consuming.

Katrin (47:03)
Yes.

Johan van de Werken (48:02)
And then all the companies that ru rely heavily on those workflows have a big problem. So and you can already see some companies that fired people are rehiring again because it turns out people are cheaper than using AI sometimes. In our specific industry, I'm not sure where we're heading for now. for the next couple of years I still think the the business logic, the business domain logic is hard to automate or to to let it to let an AI agent run all the analysis and all the and all the work we do. To do implementations, to write code, yeah, sure. you do need do don't need me for it. But to to really know your organization, that's hard to

Katrin (48:02)
It is.

Johan van de Werken (49:03)
to exclude humans from that. So

Katrin (49:08)
It is, I agree. It it it's very hard. And so you know, if we if we take like your analyst job, right? The analyst part of your of your skill set currently and the way you see it being disrupted, you know, as you said, if it's to write code or whatever, you don't need me, but you know You need me for all of these this like general like this deep knowledge and understanding of how how all of these variables weigh in real life practically in order to create something actually useful for you. Something that is going to actually help you make this better decisions and you know and and make you some money ultimately. for the analyst listening who is, you know, trying to make like a five year Five year bet is crazy, but like a one year bet on skill set. I think it's already crazy.

Johan van de Werken (49:59)
That's that's already crazy. Yeah, that's true.

Katrin (50:04)
What would you think of what are you thinking for yourself as durable and what are you really mo sort of like more and more moving away from?

Johan van de Werken (50:16)
yeah, so the durable part would be do you understand what you're talking about or what you're asking? Do you so if you ask a qu a question, the only way to to be able to check if the answer makes sense is to to be a bit more higher in your knowledge about it. Maybe you don't know that specific answer, but you need to to be able to judge if that's if it makes sense. So that's the the skill you need, I guess. What what doesn't make sense to invest your time in is learning how a user interface works. Maybe learning how to code doesn't make sense, but on the other side it will teach you stuff about cons on a conceptual level that you might profit from later on. So

Katrin (51:17)
I couldn't agree more. Use like the user interface has like we are in the world of headless. You can just use C P as long as you know what it does and you can tell it to do it, you don't need to know how the buttons work or don't work together.

Johan van de Werken (51:33)
And indeed. So if you're in the in the context of web analytics, for instance, th think about what a session means for you. We we probably know what a session is in Google Analytics, how it's defined, but do you agree with that? What is a session for you? How should a session look? Or or what is a user? What's a user? Answer that question. Because there are million different options to answer that question.

Katrin (51:33)
Yes.

Johan van de Werken (52:03)
And the the only way to to be able to talk about this if you understand everything that it that you need to know. about this world that we're working in. So what is a user? It's you can answer it from a technical point of view or from a UX point of view or there are dozens of options. So the the the simpler the question is, the harder sometimes the answer.

Katrin (52:33)
And ultimately it is a matter of decision, right? It's like the QA example, where at some point you have to decide what is ground truth. You have to at some point decide what within the multiple possibilities of what a user can be for you, what do you define as a user and why did you

Johan van de Werken (52:33)
Exactly.

Katrin (52:51)
define it this way? And this is your choice. It has gr it is grounded in the reality of your business and choices you made, and you work with that. And so that that's obviously something that has to be taught as well. And you also, I suppose, in a lot of your work, teach it to LLMs. Do you feel that now that you're teaching a lot to LLMs, is that different to you than like creating something where you teach to humans? Do you do you see similarities? Do you see differences? Do you think of the process similarly or differently?

Johan van de Werken (53:32)
So I didn't I didn't write the context files for LLMs myself, so I actually never did it. The only way but I'm doing it all the time when I'm using LLMs, right? Basically it's the same thing. You're correcting it all the time, basically. You

Katrin (53:41)
All the time.

Johan van de Werken (53:50)
give so y the first question maybe you provide context and then the correction start because

Katrin (53:50)
Yes.

Johan van de Werken (53:57)
the first answer is rarely the def defin definitive answer, right? So so in a way, maybe we are teaching agents now instead of humans. Maybe maybe I'm still an educator in that sense.

Katrin (54:12)
I find it actually really fascinating because when you start working with a new model, you you sort of, at least I, spend some time sort of learning about quote unquote the model's personality. Some things have been trained into the model, and the model is optimized for certain types of processes. also cost, right? It's it's very

Johan van de Werken (54:12)
Mm-hmm.

Katrin (54:41)
much optimized towards merging towards getting you to an answer that is satisfying you at you know at a certain level of cost. And so for that it has to make some choices and it has to make s some compromises. And different models do really handle that very differently and have different capabilities. So from that perspective in a way you sort of have to reverse engineer to a certain degree what are the model's intensive incentives and what has been trained into it so that you can do your corrections correctly and and catch catch the ways the model is always going to, you know, always going to try to sort of like turn around. And and and that's I think a learning curve every time you work with a new model. And in in a certain way, I find that that is a little similar. to teaching, you know, a human because you have to understand what a person's construction of the world is and what

Johan van de Werken (54:41)
Right.

Katrin (55:38)
prior knowledge they have, what they're going to accept, refuse, understand easily, not easily, and what type of sort of level of language and complexity you have to bring the concepts at so that they they stick. Because you you you want to teach, but you want things to stick obviously. so in in in a certain way I find that it is to a certain degree a similar process.

Johan van de Werken (56:03)
Yeah, it is. But there is a difference that humans tend to refuse stuff a bit earlier than LLMs will. So

Katrin (56:13)
yes, that's true, but LLMs do refuse stuff too.

Johan van de Werken (56:18)
Yeah, when it comes to, you know, security and that kind of stuff. Mm.

Katrin (56:24)
complexity. It's it's it's it's it's really interesting at this point, to see some of the behaviors of some of the models, you know, when you when you work with very long, very complex tasks refusing to do certain things, like literally refusing to do it. And then and then you you you try to sort of like get in and you like, no, this is a job, this is what you need to do. And then they find one way to refuse it and another way to refuse it. It's really fun it's really, really interesting.

Johan van de Werken (56:24)
stubborn agents

Katrin (56:24)
Yes.

Johan van de Werken (56:57)
or stubborn yeah models.

Katrin (57:01)
Well yes, and then you s like it's really about understanding and reverse engineering what has been trained into them, like why do they refuse this? What what is what is the reason behind it? So that you can then turn your instructions into into I don't know, chunking the the the the task in a in a in a different way. you know,

Johan van de Werken (57:01)
Right. Yeah.

Katrin (57:18)
it's it's one of the things that for example for the the the QA process, it's it's one of the things we have to work around because you're asking it to do a lot of screenshots and a lot of comparisons. So it's it it's it's you know on the large dashboard, this is it's a lot, right? It can be it can be and can

Johan van de Werken (57:18)
It is.

Katrin (57:35)
be it can be truly a lot of things. so yes, they can both refuse the tasks, but you're right, I agree. The humans tend to refuse them them often a little earlier.

Johan van de Werken (57:35)
Yeah.

Katrin (57:50)
Well, thank you, Johan. This was really great. It was a pleasure having you. so shameless plug time. where should people find your writing, your courses, your freelance sort of how do they contact you, how do people work with you?

Johan van de Werken (58:08)
I also have my own company, selectstar.nl. I'm based in the Netherlands. You can approach me there. You can still get the video course from teamsimmer.com, I believe. And of course, if you're interested in GA4Dataform, please at least install the free version, try it at ga4dataform.com.

Katrin (58:28)
Great, we'll put all the links in the show notes, GA4Dataform, the QA repository, all of it. and Johan, thank you again. That's it for episode eighteen of Knowledge Distillation. If today's conversation

Johan van de Werken (58:28)
Thanks.

Katrin (58:43)
made you want to experiment with AI for analytics, visit us as at ask-y.ai, download the plugin for QA for QA, give us some feedback, and try Prism. Thank you for listening and remember bots won't win, AI analysts will.

← Previous Episode Kelly Wortham (Forward Digital, Test & Learn Community, Experimentation Island) on Why Agentic Commerce Breaks Experimentation, Measuring What Visitors No Longer Do, and Optimizing for Machines