Connecting Braze to Snowflake, Databricks and BigQuery with Cloud Data Ingestion
By the Stitch team · · 9 min read
Key takeaways
- CDI supports five warehouses (Snowflake, Databricks, BigQuery, Redshift, Microsoft Fabric) plus S3, Azure Blob Storage and Google Cloud Storage.
- Syncs are driven by an
UPDATED_ATtimestamp: get that column right and most problems go away. - Recurring syncs run every 15 minutes by default, down to every 5 minutes on request, and you can trigger syncs on demand.
- CDI is a strong default when your warehouse is the source of truth. A reverse ETL tool earns its place when you need heavier transformation, orchestration or many destinations.
- Design nested objects carefully. Every key in a nested object can count as a data point on legacy contracts, and payloads are capped.
Which warehouses does Braze Cloud Data Ingestion support?
According to Braze's CDI documentation, CDI connects to:
- Cloud data warehouses: Amazon Redshift, Databricks, Google BigQuery, Microsoft Fabric and Snowflake
- Cloud file storage: Amazon S3, Azure Blob Storage and Google Cloud Storage
Braze says CDI is available in all Braze regions and any Braze region can connect to any source data region, which matters if you're on AU-01 and your warehouse sits elsewhere.
Each warehouse has its own connection set-up:
- Snowflake: key-pair authentication using a public key Braze provides, with a unique Snowflake user per Braze workspace. Auto-resume should be on for the warehouse.
- Databricks: an OAuth service principal or a personal access token. Classic and Pro SQL warehouses can add two to five minutes of warm-up; serverless reduces this but "may result in slightly higher integration costs".
- BigQuery: a service account, via Workload Identity Federation or a service account key, with access granted to the relevant datasets.
If your warehouse uses network policies, allowlist the Braze IPs for your instance.
How does a CDI sync work?
You create a table or view in your warehouse that Braze reads from. For the standard table set-up, each row needs:
UPDATED_AT: when the row was added or last changed- An identifier:
EXTERNAL_ID,ALIAS_NAMEwithALIAS_LABEL,BRAZE_ID,EMAILorPHONE, with one identifier type per row PAYLOAD: a JSON string of the fields to sync
On each run, Braze syncs rows where UPDATED_AT is later than the last value it synced. That makes syncs incremental, but only if your timestamps are honest.
Payload limits are worth knowing upfront. Braze says each row can carry a JSON object with up to 250 attributes, and payloads over 1 MB are rejected. By default, each run can sync up to 500 million rows.
What types of data can you sync?
CDI covers user and non-user data:
- User attributes, including nested custom attributes, arrays of objects and subscription statuses
- Custom events, with
namerequired andtime,app_idandpropertiesoptional - Purchase events, with
product_id,currencyandpricerequired - Subscription states, as
subscription_group_idandsubscription_statepairs - Catalog items, for product, content or offer catalogs
- User deletion requests
Field-level requirements for events, purchases and subscriptions are on the table set-up page. If an event row has no time, Braze uses UPDATED_AT as the event time. Braze's best practices say you can delete users by external ID, user alias or Braze ID.
Braze also offers zero-copy options. According to its comparison of ingestion options, CDI Segments and CDI Canvas triggers keep warehouse data in place rather than writing it to profiles. Those queries run in your warehouse, so you pay the compute, but Braze doesn't log data points for them.
How often can CDI sync?
Braze says recurring syncs can run as often as every 5 minutes or as rarely as once per month. The dashboard default minimum is 15 minutes, and 5-minute syncs need Braze Support or your customer success manager.
You can also trigger a sync through the API once a warehouse job finishes, which is often the cleanest pattern: your dbt or orchestration run completes, then Braze pulls the new rows. Only one sync can run per integration at a time.
Two other things to plan for:
- Rate limits: CDI shares the rate limit with the Braze API, so heavy syncs can compete with real-time
/users/tracktraffic. - Real-time needs: if something must hit Braze within seconds, such as a booking or a password reset, send it via the SDK or REST API. Braze describes
/users/trackas near real time.
How do you get Braze data back into the warehouse?
The return path matters as much as ingestion. You need engagement data in the warehouse to measure lift, build attribution and feed models.
- Snowflake Data Sharing: Braze shares a read-only database straight into your Snowflake account. What you get depends on your Data Distribution entitlement: engagement events, user behaviour events, or Profile 360 with profile and attribute changes. Shared data uses no storage in your account, so you pay only for compute when you query it.
- Currents: Braze's real-time stream of engagement events, delivered in Avro format. The available Currents partners for storage are Amazon S3, Google Cloud Storage and Azure Blob Storage, so for Databricks or BigQuery you'd typically land files in storage and load from there. Braze says a Currents connector is included in many pro and enterprise-level packages.
- Newer options: In June 2026 Braze announced User Profile Streaming and a Delta Sharing integration (beta) for Databricks that sends engagement signals back to the warehouse.
When should you use CDI instead of a reverse ETL tool like Hightouch or Census?
Braze itself says many reverse ETL use cases can also be handled with native CDI, and that the choice depends on transformation, orchestration and data ownership.
Here's how we'd frame it:
Use CDI when:
- Your warehouse is the source of truth and your data team can shape tables for Braze
- Braze is the main or only destination
- You want fewer tools and contracts to manage
Consider a reverse ETL tool like Hightouch or Census when:
- You need to send the same audiences to Braze and to ad platforms, CRM or support tools
- Marketers want a visual audience builder over warehouse models
- You want change detection handled for you. Hightouch, for example, says it only syncs data that needs updating and sends only changed columns, though it notes this can be slower than overwriting
Plenty of stacks use both: CDI for core profile and event data, and a reverse ETL tool for cross-channel audiences.
What are the common CDI pitfalls?
- Bad
UPDATED_ATvalues. A view usingCURRENT_TIMESTAMPre-syncs everything every run. Future timestamps make later syncs skip rows. Late rows with older timestamps never sync. Braze recommends UTC timestamps and writing rows that share a timestamp in a single transaction. - Re-syncing full profiles. Braze updates or adds fields, so you don't need to sync the whole profile each time. Send changes only.
- Data point consumption. If you're on a legacy data point contract, CDI is billed like
/users/track. Braze says that as of Spring 2025 it no longer bills new customers on data points, but older multi-year contracts may still be. Check yours. - Nested objects. On data point billing, every key in a nested object counts as a data point, and arrays of objects consume a data point for each key updated. Deep nesting also makes Liquid templates harder to maintain.
- Type drift. Keep column types consistent between syncs, cast numbers explicitly and store dates as timestamps, not strings.
What does CDI look like in a real implementation?
Stitch was lead Braze partner for Serko AI, working alongside Braze's onboarding team. Serko's data flows from its data warehouse into Braze via Cloud Data Ingestion, alongside the Braze Web SDK and a server-side integration via GTM.
Multi-leg trips were the tricky part. We co-designed a flattened active_flights array schema with Braze to work around nesting limits, then used Liquid to personalise messages with complex flight, hotel and cost objects. The lifecycle layer shipped to production in six weeks from kick-off, with booking confirmation, disruption SMS and 24-hour reminder flows active from day one.
How Stitch can help
Stitch is an independent Auckland consultancy and a Braze Alloys Solutions Partner (Orbit tier), named ANZ Rising Star of the Year at Braze's inaugural ANZ Partner Awards in 2026. We're also a Hightouch Implementation Partner and Segment Certified Partner, so we can help you pick the right ingestion pattern rather than defaulting to one tool.
We design warehouse schemas for Braze, set up CDI and the return path, and build the lifecycle programmes that use the data. Typical implementations take 8 to 12 weeks.
FAQ
Does Braze Cloud Data Ingestion support Databricks and BigQuery? Yes. Braze lists Snowflake, Databricks, Google BigQuery, Amazon Redshift and Microsoft Fabric as supported warehouses, plus S3, Azure Blob Storage and Google Cloud Storage.
How often can Braze CDI sync? Every 15 minutes by default, down to every 5 minutes if you ask Braze Support, or as rarely as monthly. You can also trigger a sync via the API after a warehouse job finishes.
What is the UPDATED_AT column for? Braze uses it to find rows added or changed since the last sync. If it's wrong, rows get skipped or re-synced, so keep it accurate and in UTC.
Can CDI delete users in Braze? Yes. Braze supports user deletion requests through CDI, using external ID, user alias or Braze ID.
Do I need Hightouch or Census if I have CDI? Not always. CDI covers many warehouse-to-Braze use cases. A reverse ETL tool helps when you need many destinations, visual audience building or managed change detection.
Sources
- https://www.braze.com/docs/user_guide/data/unification/cloud_ingestion
- https://www.braze.com/docs/user_guide/data/unification/cloud_ingestion/integrations/
- https://www.braze.com/docs/user_guide/data/unification/cloud_ingestion/table_setup/
- https://www.braze.com/docs/user_guide/data/unification/cloud_ingestion/best_practices/
- https://www.braze.com/docs/user_guide/data/unification/cloud_ingestion/faqs/
- https://www.braze.com/docs/user_guide/example_library/data/compare_data_ingestion_options
- https://www.braze.com/docs/user_guide/data/infrastructure/data_points
- https://www.braze.com/docs/user_guide/data/infrastructure/data_centers/
- https://www.braze.com/docs/partners/data_and_analytics/data_warehouses/snowflake/data_sharing
- https://www.braze.com/docs/user_guide/data/braze_currents/
- https://www.braze.com/docs/user_guide/data/distribution/braze_currents/setting_up_currents/available_partners
- https://www.braze.com/resources/articles/2026-braze-data-platform-product-launch
- https://www.braze.com/resources/articles/exploring-the-technical-side-of-ingesting-data-into-braze
- https://hightouch.com/docs/destinations/braze
Talk to a Braze partner
More on Braze
- Braze vs Customer.io: A Consultant's Comparison
- Braze data residency in Australia and New Zealand: what AU-01 means
- How Braze pricing works (and how to budget for it)
- Braze implementation checklist and timeline: a 12-week plan
- Migrating to Braze from Salesforce Marketing Cloud, Klaviyo or HubSpot
- BrazeAI Decisioning in practice: what it is and when it's worth it
- ANZ customer engagement benchmark 2026: what the Braze data says, and what to do about it
- Braze or Customer.io?
- Braze readiness scorecard