Blog | FELD M

Agentic analytics tools: What three months of testing taught us | FELD M

Written by Sandy Bekheet | Jul 21, 2026 1:14:03 PM

Picture the scene: You ask an agentic analytics tool for revenue from returning customers in 2025. You get your numbers, ask a few follow-ups, then grab a coffee and have a chat with a colleague. You come back, settle into your chair, and type "show me customer refunds." You were, of course, expecting to see refunds for this year (2026), for both new and returning customers.

One tool shows total refunds for 2025 (still referencing the previous prompt where you asked about last year's returning customers, even though 2026 was also expected).

Another keeps filtering on the returning customers you were asking about three messages ago.

The third stops and asks: Should it be split by customer type?

The same prompt, three tools, three different answers, and only one of them is what you actually wanted.

 

Testing our way through the agentic analytics tool market

That gap is why we spent the first quarter of this year testing the different agentic analytics offerings out there, and the deeper we got, the more questions we had. Our objective was to find tools we like and can recommend confidently to our clients.

The initial question was practical: which of these tools can we trust enough to recommend to a client? The second crept up on us: if an agent can go from a plain-language question to a finished dashboard, what does it mean for our own jobs as BI specialists?

The rush of agentic analytics tools had been building slowly over the past year, but it still felt like there was almost a new tool coming out every month. Our Data Product Team had already categorized some tools and tested others, but it was time to do more: define guidelines, prioritize tools, reach out for demos, and dig into some actual testing.

Today, we're sharing the perspectives and experience of two of our team's BI experts, Dima Yarmolin and Sandy Bekheet. Both questions ran through everything they tested, but each of them kept coming back to a different one:

  • For Sandy, it was trust: stress-testing the answers and catching the moments a tool would confidently get it wrong.

  • For Dima, it was the role: noticing his own habits break and rebuild as the work shifted from building every answer to preparing the ground for agents.

They agreed on a lot, and where they didn't, well, we left those moments in too!

 

 

 

Testing guidelines

With coffee, cake, and plenty of project experience, we started thinking about our concrete testing criteria. Luckily for us, we already had the perfect project to test with:  Forever Thirsty, a natural wine shop in Munich owned by our CEO, Lutz Wiechert (you should stop by when you're in Munich!). Forever Thirsty's data lives in BigQuery, with wide tables transformed and ready through dbt.

Here's exactly what we looked for in every tool:

 

Click here to view the above image full size.

Now, onto the interesting part: How did the testing go?

 

Sandy's personal experience: Establishing trust through testing

"Testing these tools wasn't about ticking feature boxes. I kept asking one question: would I trust recommending any of these tools to one of our clients? That bar changed everything about how I evaluated them, and honestly it's left me excited about how these tools will shape the future of business intelligence."

Sandy Bekheet, FELD M

We'd already started trying tools like Veezoo and ThoughtSpot when they first appeared on the market. They had a lot of promise, and comparing them back then and now, it's exciting to see how far we've come.

I assigned myself the task of defining the whole evaluation process and bringing in colleagues to test and challenge the tools together. The fun part was meeting interesting people, whether on the sales side or technical product owners, who walked us through their tool's features and the areas where their tools are especially strong.

Tools rarely won everywhere against our criteria. Some scored high on the development process or agent creation; others on visualization, speed, answer quality, or statistical models. The one thing they all had in common: large language models (LLMs).

I was impressed by how fast a tool could go from a simple question to designing a full dashboard. You could simply check the SQL, make some small adjustments, and you were good to go. This will save developers a lot of time to focus on business questions and logic, which leads us nicely to a bigger topic that Dima will dive into later in this post: how BI roles are changing.

Still, I did miss the customizability of BI tools to do things like create parameters, build filters with different types, customize tooltips, and cross-filter.

But that kind of manual customization may soon be unnecessary. Instead, you could filter further via prompting and dig into the details yourself, on top of the basic filtering some tools already provide (for example, a date selector). A tool like Veezoo, for example, allows for high customizability through its canvas feature, while Data Studio Pro can't build a dashboard via prompting.

Whether the skill of building a dashboard is becoming optional or whether it's something clients will end up missing is a question I don't think the industry has answered yet. It's the thing I'll be watching most closely.

What was challenging is that you have to keep your eyes wide open all the time while testing. You're basically challenging the AI: asking trick questions, running hallucination tests, rephrasing, simulating what a business user would ask on a normal business day. The customer returns scenario we mentioned earlier in this post wasn't made up. It's exactly what happens when you stop feeding the tool perfect prompts and start using it the way a busy person actually would. How tools handle this varies, and that's a design problem, not a user error. Of course, the question could be more precise, but with the speed of daily work, it won't be. A good tool either asks or makes its assumptions visible.

Another scenario worried me more. You could ask for last month's marketing revenue without knowing that one of the most important data sources has no data for last month due to an issue. If the tool doesn't flag it, it gives you a misleading response, and you'd never know to question it.

Other challenges included speed, limited chart types or visualization features (like filtering within a dashboard or parameter creation), and no row-level security. And even when I was enthusiastic about a certain tool during the demo, it didn't mean it would live up to the hype. Unfortunately, some vendors wouldn't provide us with a trial unless we already had a potential client project lined up. That's understandable, but how can we recommend a tool to our clients that we've never tested?

A very important last point I won't go through in detail here is about semantic layers. It's the layer that lets the AI agent correctly map your question to your data. It's where you define your metrics, business terms, and how everything relates. The more you define and curate it, the higher the quality of the answers you get back.

In my testing, the pattern was clear: tools without a strong semantic layer delivered terrible results. The only exception was when there were so few data sources in shape (e.g., cleaned, aggregated, joined) that the metadata alone was enough to carry it, like with Data Studio Pro.

 

Dima's personal experience: Learning to work differently

"The goal was always to find tools we could confidently recommend to our clients. But somewhere along the way, I also started asking what my own job will look like in a few years. That turned out to be just as interesting a question."

Dima Yarmolin, FELD M

I came into this with a classic BI mindset. Years of Tableau, Looker Studio, and Power BI train you in a certain way: you model the data, you build the views, you place every filter and calculated field yourself. So in the first testing sessions, I caught myself working exactly like that: looking for the field list, planning the layout in my head. And it just doesn't apply here. These tools run in completely different conditions. You're no longer building every answer from scratch — you're preparing the foundation so that the LLMs and AI-powered agents can build it well. That flips the whole workflow, and changes what a BI specialist even is.

Once I accepted that, I noticed that my habits changed when evaluating the tools. A big part of the job now is judging answers instead of building every one of them: asking the same question in a few different ways, checking over the generated SQL, staying curious about how the tool got to its result. And I don't see that as distrust. It's the same instinct that makes you test your own dashboard before you ship it, just aimed at something you didn't build yourself.

The building side was where things became genuinely fun. When the semantic layer is set up right, the work starts with a conversation instead of a blank canvas. You explore by asking questions, one follow-up after another, and you've formed a picture of the answer before any dashboard even exists. From there, I moved into building, turning those answers into a dashboard — or in Veezoo's case, a canvas — for Forever Thirsty and styling it to match their website, so it felt like their product and not a generic report. Building a whole dashboard just by prompting is a genuinely fun and new way to work.

Still, just like Sandy, I found myself missing the level of customizability that Tableau or Power BI give you, where you shape every filter, tooltip, and detail exactly how you want them. When you're shipping a client-facing page that has to look and behave just right, that kind of finer control still has an edge that these tools aren't able to fully replicate. Sandy and I land a little differently here, though; she's more ready to let that go than I am, and she may well be right that we'll miss it less as the tools mature.

 

Where we land

Where we fully agree is on what comes next, and it's the part that really changed my thinking: a dashboard doesn't have to be a static end product anymore. You can put agents behind it, let them run parts of the analysis, and deliver the findings in dashboard form. And that shift isn't tied to one tool or one shape.

The analytics platform, Hex, showed me the same idea from a more notebook-style angle, where agents take over the deeper analysis steps, and you shape what comes back into something repeatable. Different tools, same direction: the analysis itself is starting to run on its own, and our job is to design the "how".

Every test, every failed prompt, every single insight we gained became shared knowledge in our Data Product Team, and that's worth more than any single tool recommendation. And that knowledge doesn't stay internal for long: the moment a client has a question one of us has already worked through, the answer comes back sharper and the conversation moves to decisions instead of mechanics.

It also gave me a clearer picture of where the BI role is going: less manual building, more shaping context, judging quality, and designing how people and agents work together. The funny thing is that the old fundamentals — clean data modeling and business logic — matter more than ever, because that's exactly what the agents depend on.

And looking at where the market stands, I think our timing is good: these tools are ready enough to learn from, and the field is young enough that what we build today still puts us ahead.

 

Summary

In the end, we tested more than 15 tools and narrowed our favorites down to around 5. Here's an overview of those 15 tools:

 

 

And here's what we believe after all of it: these tools are ready to change how BI work gets done, but none of them are ready to be trusted unsupervised. What sets apart the best ones is that they don't just provide answers fast — they also tell you when something looks off.

Choosing the right tool for your company still depends on so many factors: where you want your semantic layer defined, how complex you want your analysis to be, how technical your data team is, whether your business users are ready to know what to look for when working with agents, and how much you're willing to invest.

Sandy has also put together a list of the questions our clients ask most often about AI adoption in her post Do you really need an agentic analytics tool?

Building AI skills across an organization is easier said than done. Our AI literacy canvas lays the whole picture on a single page: who needs which skills, for which use cases, in what format, and how to measure whether it's working. If you'd like support putting it into practice, our AI literacy training covers everything from a two-hour introduction to a two-day deep dive.