![[Presentation Slides] I gave a presentation titled "Let's Have Snowflake's AI Analyze Data While Breaking Data Quality" #devio2026](https://devio2024-media.developers.io/image/upload/f_auto,q_auto,w_3840/v1790758605/user-gen-eyecatch/sidmcbh1yxyyvjoejhsj.png)
[Presentation Slides] I gave a presentation titled "Let's Have Snowflake's AI Analyze Data While Breaking Data Quality" #devio2026
This page has been translated by machine translation. View original
This is Kawanago from the Data Business Division.
Recently, there was an event called DevelopersIO 2026 Osaka at Classmethod Osaka.
It is an offline event where people discuss a wide range of IT technologies.
I presented there under the title "Letting Snowflake's AI Analyze Data While Breaking Data Quality",
so I'd like to briefly introduce the content of my presentation here as well.
Why I Decided to Talk About This
This DevelopersIO Osaka event is held around September every year,
and topics and content for presentations are募集 internally around July.
About a month before the call for presentation topics, I had given a talk on data management,
and remembering that I had received a comment saying "I'd love to hear more concrete details",
I decided to talk about data quality from a slightly more implementation-oriented perspective.
I also decided to tie it in with AI because I was personally interested in it.
Lately, data analysis using AI agents has been spreading rapidly,
and I was curious about how much data quality affects that.
So the content became about touching on the impact of data quality on AI-driven analysis,
while also trying out data quality in Snowflake, which I have been working with frequently lately.
The Impact of Data Quality on AI
The verification is based on a configuration image like the following.
On a general medallion architecture, a Semantic View is created in the GOLD layer.
The configuration is such that Cortex Analyst generates SQL by referencing that View.

The target of the analysis is an order table like this.

When I asked about August sales in a normal state, I got an answer of 915,900 yen.

From this state, I break the data quality in various patterns.
In the following pattern, the August data is simply duplicated.

When I asked the same question in this state, the amount came out to twice that of the normal state.

I also verified about two other error cases, but the results were similar.

Even if the SQL being executed is correct, if the quality of the data itself is poor, the results will not follow.
If AI generates incorrect results, there is a risk of giving users incorrect insights.
However, since I intentionally selected such cases for this verification,
there were also cases where it properly noticed and pointed out the issues depending on the content.
Snowflake's Semantic View has descriptions of tables, columns, and relationships,
but it does not hold information about data quality, so a separate mechanism is needed to protect data quality.

About Snowflake's Data Quality-Related Features
Snowflake has DMF (Data Metric Functions) as a feature for protecting data quality.
It allows you to set functions that return data quality values for tables and columns,
and can be executed on a schedule or when data changes are detected.
Additionally, by setting Expectations together,
you can create a mechanism to send notifications when values fall below a preset threshold.

Many DMFs are available out of the box and are easy to configure,
and the fact that they are natively integrated with Horizon Catalog is extremely convenient.

Also, although it is in preview, an anomaly detection feature has been implemented as well.
The sensitivity can be selected from 3 levels, so it seems capable of handling a certain range of use cases.
This is definitely a feature I'd want to use in cases where it's difficult to determine normal thresholds in advance.

On How to Use It Alongside dbt
There was something I kept thinking about while putting this material together.
That is: "Isn't dbt's test feature good enough for this?"
In Snowflake development environments, dbt is often used as a data transformation tool,
and I think it's standard practice to have detailed tests on each model during implementation.
If you are using dbt, I think basic data quality can be ensured with dbt's test feature.
Also, since dbt allows code-based management, from an engineer's perspective, dbt seems to have greater advantages.
However, I think the biggest benefit of using Snowflake's data quality features is
that they are natively integrated with Horizon Catalog.

With dbt core, test results are not displayed in a GUI anywhere,
so users who use the data cannot tell what the state of the data quality is.
Since Snowflake's DMF and anomaly detection results are displayed on the catalog,
the state of quality can be made visible to data users.
In organizations where data quality targets have been agreed upon with data users,
I think there is merit in using it alongside dbt for that reason alone.
In Closing
AI-driven data analysis and utilization is extremely cutting-edge and impactful,
and I think many organizations are positively inclined to pursue it as a means of improving operational efficiency.
On the other hand, data quality management is very unglamorous and tedious work,
and there are many cases where it has not kept pace with data utilization.
Data quality is the foundation and prerequisite for analysis,
so regardless of the tools used, it is important to first build a mechanism for understanding the state of data quality.
DMF is very easy to configure, so please give it a try.
Thank you to everyone who participated in this event.
I hope the presentation content and this article will be of some help to you.
Thank you for reading this article to the end.



