Yeedu Hits $0.53/TB in TPC-DS Benchmark
Yeedu
HomeBlogsDatabricks and Yeedu: How They Work Together
Blog

Databricks and Yeedu: How They Work Together

Mayank MehraOctober 6, 2026
Databricks and Yeedu: How They Work Together

Cloud data costs keep rising, and for many businesses the Databricks bill, measured in Databricks Units (DBUs) plus the underlying cloud compute, is one of the biggest lines. The usual advice is to switch platforms. That means months of migration, retraining and risk. 

There is a lower-effort way. Keep your Databricks lakehouse, and run only the Spark jobs that cost you the most on Yeedu.

Why 20% of Databricks jobs drive half the DBU bill 

Say you run 1,000 production jobs on Databricks Jobs compute. They don’t all cost the same. In most environments, about 20% of jobs consume up to 50% of the budget. 

Those jobs are usually long-running PySpark, Spark SQL or Scala pipelines with heavy joins, aggregations and shuffles. That 20% is where Yeedu helps. 

Yeedu runs alongside your existing Databricks workspace, and our customer case studies show it cutting production costs by around 60%. The savings come from Turbo, Yeedu’s re-architected Spark execution engine, which runs the same jobs on less compute. 

Keep the Databricks lakehouse, offload the high-DBU jobs 

For most enterprises, the biggest barrier to cutting platform costs is the migration itself. So we don’t recommend replacing Databricks. 

The unit of adoption is a job, not a platform. Start by running your highest-cost workloads on Yeedu, and leave notebooks, dashboards, Unity Catalog and every other job where they are. 

This gives you:

  • Savings early: You start with the jobs that cost the most, so the impact shows up in the next billing cycle. 
  • Time to plan: Lower-cost jobs can follow later, on your schedule, or stay on Databricks. 
  • Confidence before you commit: You see real results on real workloads before extending the rollout. 

Notebooks, Git repos and Databricks Workflows: what happens to your code 

Every migration conversation comes down to two questions: what happens to our code, and what happens to our data? For code, your team keeps working the way it does today. 

  • Notebooks carry over: Yeedu supports an open notebook format, so notebooks import and export between Databricks and Yeedu. Yeedu notebooks run Python, PySpark, SQL, Pandas and DuckDB. 
  • Same Git repositories: If your Databricks code is synced to GitHub through Git folders, your team connects the same repositories in Yeedu. Version control stays in one place. 
  • A migration utility does the heavy lifting: It moves your code and pipelines from Databricks, recreates DBFS mounts and handles custom dbutils calls such as dbutils.fs and dbutils.secrets. 
  • Workflows keep their schedules: Databricks Workflows (now Lakeflow Jobs) become Yeedu pipelines with the same schedules and task order.

Unity Catalog, Delta Lake and governance: your data stays put 

Your data stays in Databricks Unity Catalog. Yeedu connects to it as a foreign catalog through its Metastore / BYOCatalog layer and works with the data where it already lives. There is no data migration project and no duplicated metadata. 

  • Same access controls: Each user adds their Databricks personal access token (PAT) to their Yeedu user secrets. The Unity Catalog grants and role-based access you already defined apply as is, with no code changes. 
  • Read and write: Yeedu reads from and writes to Unity Catalog tables. A few write exceptions are listed in the Yeedu Databricks migration docs. 
  • Governance stays intact: Unity Catalog lineage and audit records for Yeedu’s reads and writes still appear in Databricks. 
  • Both platforms share Delta tables: A Databricks job and a Yeedu job can safely read and write the same Delta Lake tables, with the Delta transaction log keeping writes consistent. 
  • Same connections to other sources: Secrets for Amazon S3, ADLS Gen2, JDBC databases, Kafka and Pub/Sub are set up once in Yeedu, and existing dbutils.secrets.get calls work unchanged. 

Yeedu also reads Hive Metastore and AWS Glue, and works with Iceberg as well as Delta tables, for teams with more than one catalog.

Which Databricks workloads run on Yeedu: batch, streaming, Auto Loader, MLflow and BI 

Yeedu can run almost 95% of typical Databricks jobs and workloads, across data engineering, data science and analytics.

Databricks Workload How It Runs on Yeedu
Batch ETL and ELT jobs PySpark, Spark SQL and Scala jobs run as they are, with no code changes
Event-driven jobs (Kafka, Pub/Sub) The same Spark Structured Streaming code that listens to your event platforms runs on Yeedu, with no code changes
Incremental file loading (Auto Loader) Moves with no rewrite, on S3, GCS, ADLS Gen2 and other S3A-compatible storage; customer results show more than 50% cost reduction
ML inference Yeedu Functions deploy models or Python programs as HTTPS REST endpoints with bearer-token auth, scaling automatically with request volume
Model training Open, MLflow-compatible workloads run inside Yeedu
BI dashboards and SQL analytics Yeedu’s Thrift server connects over JDBC and ODBC to tools such as Tableau and Power BI

Batch ETL and ELT jobs

How It Runs on Yeedu PySpark, Spark SQL and Scala jobs run as they are, with no code changes

Event-driven jobs (Kafka, Pub/Sub)

How It Runs on Yeedu The same Spark Structured Streaming code that listens to your event platforms runs on Yeedu, with no code changes

Incremental file loading (Auto Loader)

How It Runs on Yeedu Moves with no rewrite, on S3, GCS, ADLS Gen2 and other S3A-compatible storage; customer results show more than 50% cost reduction

ML inference

How It Runs on Yeedu Yeedu Functions deploy models or Python programs as HTTPS REST endpoints with bearer-token auth, scaling automatically with request volume

Model training

How It Runs on Yeedu Open, MLflow-compatible workloads run inside Yeedu

BI dashboards and SQL analytics

How It Runs on Yeedu Yeedu’s Thrift server connects over JDBC and ODBC to tools such as Tableau and Power BI

For dashboards, Yeedu Warmstart has a cluster ready in under 20 seconds, and Turbo runs the same SQL queries faster on less compute. 

One current gap: Yeedu doesn’t offer no-code connectors like Databricks does. It supports all open-source PySpark connectors, so those integrations can be written as code and run on Yeedu.

Inside the Turbo engine: how Yeedu runs the same Spark jobs on less compute 

Turbo is a C++ execution layer underneath the Spark API. It swaps the JVM execution path for a vectorized, SIMD-accelerated columnar runtime, and keeps hot data in CPU caches (L2 and L3) rather than on the JVM heap. It also rewrites query plans before execution. 

On CPU-bound work, meaning joins, aggregations, multi-stage transforms and ML feature preparation, Yeedu reports 4 to 10 times faster execution and 60 to 80% lower compute cost. Yeedu puts that slice at 30 to 40% of a typical workload mix. 

For I/O-bound jobs, Smart Scheduling detects idle CPU windows while tasks wait on reads and writes, and packs other tasks into them. Yeedu reports 2 to 4 times higher cluster efficiency for ingestion, ELT and streaming jobs. 

Around the engine, Yeedu jobs get automatic cluster start and stop, autoscaling, spark-submit multiplexing and real-time cost per job. Runtimes cover x86, ARM and GPU instances, and several Spark versions can run side by side. 

VPC deployment, golden images and compliance 

Yeedu deploys entirely inside your own cloud account and VPC, on AWS, Azure, GCP or Oracle Cloud (OCI), or on your own servers. No data leaves your network, and the deployment follows your existing enterprise architecture. 

The Spark compute Yeedu launches can use your own golden image as its base image, so your VM hardening and deployment standards stay in place. Yeedu lists ISO 27001, SOC 2, HIPAA and GDPR compliance. 

Monitoring plugs into the tools you already use, including Grafana, Amazon CloudWatch and Splunk. 

Airflow and Prefect failover to Databricks for SLA-critical jobs 

Some jobs have SLAs you can’t miss. For those, Databricks can act as your fallback. 

Orchestrators such as Apache Airflow and Prefect can submit and schedule jobs on Yeedu. In Airflow, you can configure failover so that a job that fails on Yeedu reruns on Databricks. 

A common pattern is a Yeedu task followed by a Databricks task, using the Airflow Databricks provider’s DatabricksRunNowOperator, with trigger_rule="one_failed". The Databricks run only starts if the Yeedu run fails.

Find your costliest jobs with Databricks system tables 

Databricks already records what each job costs. The system.billing.usage system table, available through Unity Catalog, shows DBU usage per job, so your team can list the jobs that cost the most. Those are your first candidates to run on Yeedu. 

Hand this query to your data team. It joins system.billing.usage with system.billing.list_prices and system.lakeflow.jobs, and lists your top 50 jobs by estimated cost over the last 30 days.

WITH job_names AS (
SELECT workspace_id, job_id, name
FROM system.lakeflow.jobs
QUALIFY ROW_NUMBER() OVER (PARTITION BY workspace_id, job_id ORDER BY change_time DESC) = 1
)
SELECT
u.workspace_id,
u.usage_metadata.job_id AS job_id,
j.name AS job_name,
SUM(u.usage_quantity) AS dbus,
SUM(u.usage_quantity * p.pricing.default) AS est_list_cost_usd
FROM system.billing.usage u
JOIN system.billing.list_prices p
ON u.sku_name = p.sku_name
AND u.usage_start_time >= p.price_start_time
AND (p.price_end_time IS NULL OR u.usage_start_time < p.price_end_time)
LEFT JOIN job_names j
ON u.workspace_id = j.workspace_id AND u.usage_metadata.job_id = j.job_id
WHERE u.usage_metadata.job_id IS NOT NULL
AND u.usage_date >= current_date() - INTERVAL 30 DAYS
GROUP BY ALL
ORDER BY est_list_cost_usd DESC
LIMIT 50;

The estimate uses list prices, so it ignores any negotiated discount. The ranking still holds. The query can also be filtered by custom_tags or workspace, to see which teams’ jobs cost the most. 

The bottom line 

Cutting your data platform bill doesn’t require a risky platform switch. Keep Databricks, run the 20% of jobs that drive your DBU spend on Yeedu, and keep your code, Unity Catalog data and governance where they are. 

Customer case studies show savings of around 60%. You can grow from there at your own pace.

Frequently Asked Questions 

Can Yeedu run existing Databricks Spark workloads without requiring pipeline rewrites? 

Yes, for most workloads. Batch PySpark, Spark SQL and Scala jobs, Structured Streaming jobs on Kafka or Pub/Sub, and Auto Loader pipelines run with no code changes. 

A migration utility moves notebooks and pipelines, recreates DBFS mounts, handles dbutils calls and turns Databricks Workflows into Yeedu pipelines with the same schedules. The main exception is Databricks’ no-code connectors, which need to be written as PySpark connector code. 

Can Yeedu help reduce Databricks costs while keeping existing Spark workloads? 

Yes. You keep Databricks and run only the highest-DBU jobs on Yeedu, which usually means the 20% of jobs behind about half the bill. 

Those jobs keep their code, their Unity Catalog tables and their governance. Yeedu’s customer case studies show production costs falling by around 60%, and Yeedu is licensed at a fixed annual price with unlimited usage, so the Yeedu side of the bill doesn’t grow with DBU-style consumption. 

Can Databricks and Yeedu be used together? 

Yes, and that’s the recommended setup. Yeedu reads and writes Unity Catalog tables as a foreign catalog, respects your existing grants through each user’s personal access token, and records lineage and audit events back in Databricks. 

Both platforms can read and write the same Delta tables. Airflow or Prefect can orchestrate jobs across both, with Databricks as the failover target for SLA-critical runs. 

How can Yeedu improve the performance of Databricks Spark workloads? 

Through its Turbo engine, a C++ vectorized, SIMD-accelerated columnar runtime underneath the Spark API. On CPU-bound work such as joins, aggregations and multi-stage transforms, Yeedu reports 4 to 10 times faster execution. 

For I/O-bound ingestion and ELT jobs, Smart Scheduling fills idle CPU time during reads and writes, with a reported 2 to 4 times higher cluster efficiency. For BI, Warmstart has a cluster ready in under 20 seconds.  

  

Join our Insider Circle

Get exclusive content crafted for engineers, architects, and data leaders building the next generation of platforms.

No spam. Just high-value intel.