Early this year, Spark Declarative Pipelines (SDP) was announced, which has made it dramatically easier to build robust Spark pipelines using a framework that abstracts away orchestration and complexity. The SDP declarative framework extends beyond individual queries to enable a mix of batch and streaming pipelines, keeping multiple datasets fresh.
In this session, Sandy Ryza (Databricks) shares a broader vision for the future of Spark Declarative Pipelines — one that opens the door to a new level of openness, standardization, and community momentum.
Key takeaways:
✅ Core concepts behind Spark Declarative Pipelines;
✅ Where the architecture is headed; and
✅ What this shift means for both existing users and Spark engineers building procedural code.
This talk was part of the “Open Lakehouse + AI Mini Summit,” hosted at the Databricks Mountain View office on November 13, 2025.
00:00 – Intro
02:00 – What “declarative” means; what is a pipeline
04:30 – Why non‑declarative pipelines are hard (ordering, parallelism)
08:00 – Errors, retries, and discrete vs continuous execution
11:30 – Orchestrators (Airflow) pros/cons and limits
14:30 – Introducing Spark Declarative Pipelines (SDP)
17:30 – Core APIs: streaming tables, materialized views, SQL/Python, runner
22:30 – Modes, dependency inference, validation; SDP vs dbt vs Airflow; availability
Nesta página do site você pode assistir ao vídeo on-line Declarative Pipelines: What’s Next for Apache Spark duração hora minuto segundo em boa qualidade , que foi baixado pelo usuário Open Lakehouse + AI 18 Novembro 2025, compartilhe o link com seus amigos e conhecidos, no youtube este vídeo já foi visto 1,925 vezes e gostou 51 espectadores. Boa visualização!