![[Presentation Material] I spoke at DevelopersIO 2026 Osaka under the title "Building an Open Data Platform with Snowflake Horizon Catalog and Apache Iceberg" #devio2026](https://devio2024-media.developers.io/image/upload/f_auto,q_auto,w_3840/v1790660895/user-gen-eyecatch/qdfyhcqs6waeeitcshep.png)
[Presentation Material] I spoke at DevelopersIO 2026 Osaka under the title "Building an Open Data Platform with Snowflake Horizon Catalog and Apache Iceberg" #devio2026
This page has been translated by machine translation. View original
Hello, I'm Kitagawa from the Data Business Division.
I presented at DevelopersIO 2026 Osaka, so I'm leaving a record of it on the blog as well.
Slides
Here's a brief introduction of the highlights below.

During the presentation I had "I like mountain climbing," but it didn't feel quite right so I changed it to "I like mountains."
My biggest regret of the day was failing to say "I can't answer everything, but I'll answer what I know" when Snowflake said "feel free to ask me anything."

The target audience I had in mind is people who have used Snowflake a little but haven't yet gotten their hands on Iceberg. Despite what the grandiose writing might suggest, I didn't talk all that much about building data platforms themselves. I showed that Iceberg tables can be read from and written to by multiple engines, along with a demo of that.
What is an Open Data Platform?

Traditionally, if you wanted to use data stored in Snowflake with Spark or Trino, for example, you would unload the data to S3 or similar storage to handle it. This creates issues with data freshness and governance. With an open data platform, the same data files can be referenced from various query engines.
Using Horizon Catalog and Iceberg can resolve these challenges.
What is Apache Iceberg?

Iceberg is a table format. While holding files on object storage such as S3, it enables ACID transactions and extremely fast queries. It also has many other features.

The aforementioned features are realized through a metadata layer with a hierarchical structure.

Also, because the Iceberg REST Catalog API is standardized, query engines can access data regardless of the service.
What is Snowflake Horizon Catalog?

Horizon Catalog is an integrated catalog that governs and protects data both inside and outside Snowflake, and connects compute resources and AI models to data. It is built on Apache Polaris.
Through Horizon Catalog, roles and policies can be applied to queries from external engines as well.
Also, Iceberg's vended credentials allow data usage to be permitted without granting external engines access permissions to S3 or similar storage.
Trying Iceberg in Snowflake

Currently, there are three main ways to use Iceberg tables in Snowflake.
- Using Snowflake Horizon Catalog with storage held outside Snowflake
- Using Snowflake Horizon Catalog with storage held inside Snowflake
- Using an external catalog
Mutual use between Snowflake and external engines is possible with any of these methods. Using Snowflake Horizon Catalog has the advantage that Snowflake handles Iceberg table maintenance such as compaction.

In the demo, I confirmed a series of steps: creating an Iceberg table in Snowflake, writing to it with Amazon Athena's PySpark engine, and referencing it with DuckDB.

Creating a table in Snowflake

Insert with Amazon Athena

SELECT with DuckDB
For Amazon Athena and DuckDB, dedicated service users were prepared and authentication is done via PAT. I was successfully able to verify referencing from multiple engines.
Reference blogs
Summary

After the Presentation
Feel free to consult me about anything related to Snowflake. I look forward to hearing from you.
