Introducing the New Python Data Source API for Apache Spark™

Databricks July 23, 2024
Video Thumbnail
Databricks Logo

Databricks

@databricks

About

Databricks is the Data and AI company. More than 20,000 organizations worldwide — including adidas, AT&T, Bayer, Block, Mastercard, Rivian, Unilever, and over 60% of the Fortune 500 — rely on Databricks to build and scale data and AI apps, analytics and agents. Headquartered in San Francisco with 30+ offices around the globe, Databricks offers a unified Data Intelligence Platform that includes Agent Bricks, Genie, Lakebase, Lakeflow, Lakehouse, and Unity Catalog.

Video Description

The introduction of the Python Data Source API for Apache Spark™ marks a significant advancement in making big data processing more accessible to Python developers. Traditionally, integrating custom data sources into Spark required understanding Scala, posing a challenge for the vast Python community. Our new API simplifies this process, allowing developers to implement custom data sources directly in Python without the complexities of existing APIs. This session will outline the API's key features, including simplified operations for reading and writing data, and its benefits to Python developers. We aim to open up Spark to more Python developers, making the big data ecosystem more inclusive and user-friendly. We will also invite one of the Databricks customers to co-present this talk. Talk By: Allison Wang, Sr. Software Engineer, Databricks ; Ryan Nienhuis, Sr. Staff Product Manager, Databricks Here’s more to explore: Big Book of Data Engineering: 2nd Edition: https://dbricks.co/3XpPgNV The Data Team's Guide to the Databricks Lakehouse Platform: https://dbricks.co/46nuDpI Connect with us: Website: https://databricks.com Twitter: https://twitter.com/databricks LinkedIn: https://www.linkedin.com/company/data… Instagram: https://www.instagram.com/databricksinc Facebook: https://www.facebook.com/databricksinc

You May Also Like