- English
- English
Appearance
Imagine you work at an e-commerce company. Data lives in many places: transactions in the app database, user clicks in event tracking, payments from a payment gateway, product data from inventory.
Then problems start:
Medallion Architecture is used to solve problems like these.
Medallion Architecture is a design pattern for organizing data flow in three layers. The names follow medal quality: Bronze, Silver, and Gold.
The core idea is simple: the business should not use raw data directly. Data is cleaned and improved in steps. At the end you have one clean source of truth that is ready to use.
Databricks popularized this idea. Many teams now use it, and it is not tied to one platform.
Bronze stores data exactly as it arrives from the source. No changes. What you ingest is what you keep.
This layer is insurance. If a pipeline above Bronze fails, you do not need to ask another team for the data again, and you do not need to extract again from the production database. You replay from Bronze: rerun the transforms on data that is already stored.
Terms you will hear at work
Some companies call this layer Landing or Staging. If a teammate says "data staging area", they often mean the same thing as Bronze or the Raw layer.
This is where cleaning happens. Raw data is transformed so it is structured, clean, and consistent.
Common work in Silver:
IDR, Rp, and rupiah become one format; timestamps are aligned to UTC.@ or a negative age.Silver is the single source of truth for detailed data: one agreed version that people treat as the reference. If the debate is "which number is correct", look at Silver.
Terms you will hear at work
This cleaning is often called cleansing, conforming (making formats match across sources), or harmonization. The layer is sometimes called the Cleansed layer.
Gold is the business-level layer. Data is aggregated and modeled for consumers (dashboards, reports, machine learning).
Example Gold tables, named after the business question:
daily_revenue_per_regioncustomer_lifetime_valuemonthly_active_usersSome teams model Gold as a star schema: fact tables for events or transactions (fact_orders), dimension tables for context such as customers or products (dim_customer). That is a pattern for how tables relate. It does not replace the descriptive names above. The point is that Gold is shaped so it is easy to consume.
Terms you will hear at work
This layer is often called the Curated layer or Mart layer. A data mart is a set of tables for one business need, for example sales.
That set usually lives in a schema or a database. Teams often use these two words for the same container, depending on the tool (Spark, Glue, and Iceberg catalogs usually say database or namespace; Postgres and Snowflake keep database and schema as separate levels). Example: schema mart_sales, table daily_revenue_per_region. In SQL it looks like mart_sales.daily_revenue_per_region. The dot separates the container and the table. It is not part of the table name.
Inside the same mart you can still have descriptive tables or fact/dimension pairs.
Follow one dataset from source to consumers.
Bronze. Transactions are extracted from the app database every hour and stored as-is. This includes failed transactions, duplicates, and columns with mixed formats.
Silver.
transaction_id.FAILED.txn_amt becomes amount, and currency format is standardized.Gold.
daily_revenue table.The result: if a data engineer finds a bug in the aggregation, they fix the logic and reprocess from Silver. They do not need to touch the source, and they do not need to panic.
| Problem without Medallion | What Medallion gives you |
|---|---|
| Raw data is overwritten during transform | Bronze is append-only, so you can always replay |
| Each team has its own "version of the truth" | Silver is the single source of truth |
| Complex transforms are mixed in one place | Each layer has one responsibility |
| A bug in one step breaks everything | You can fix one layer without touching the rest |
This is separation of concerns, like splitting presentation, business logic, and data access in software engineering. Here you split levels of data maturity instead.
The idea is not tied to one platform. On AWS, one common stack is:
bronze/, silver/, gold/, or separate buckets.On other platforms (Databricks, Snowflake, GCP, Azure) the tool names change. The pattern stays the same. That is why the concept is worth more than memorizing tools.
These show up often at work, and they are worth avoiding:
A good fit when:
A weaker fit for a small project: one source, one dashboard. The full pattern adds complexity without much value.
If you are new to data engineering, Medallion Architecture is a foundation you will see often. Many mid-size and larger data teams use a variation of this pattern.