01 / SPARK CONTROL PLANE

Change Spark with proof.

Connect AWS, run under policy and budget guardrails, diagnose failures, understand cost, and prove business results before cutover — all inside your own account.

EVALUATIONb10aac5
emr-6.15.0emr-7.5.0
“Same numbers out the other side.”
assumed, not measured
ASSERTIONS10
CHECKS DIFFERED3
FINDINGS5
NOT_EQUIVALENT
DO_NOT_CUT_OVER
ONE EVALUATION, START TO FINISH
LIVE TRACE
CHANGEEXECUTECOMPAREDO_NOT_CUT_OVER
JOB RUN 00g7nppt0p4fu80b

02 / THE PROBLEM

A runtime upgrade that finishes is not a runtime upgrade that worked.

TODAY
Spark 3.4Spark 3.5
“All 47 tasks succeeded. Row counts match.”
EVIDENCE / TASK EXIT STATES · ROW COUNTS
DECISION BASIS / NOTHING COMPARED TO THE OLD RUNTIME
WITH SPARKPILOT
Spark 3.4EXECUTECOMPARESpark 3.5
DO_NOT_CUT_OVER
EVIDENCE / 10 ASSERTIONS · 3 CHECKS DIFFERED · 5 FINDINGS
DECISION BASIS / RECOMPUTABLE TRACE

Row counts match while a decimal(38,6) → decimal(38,1) cast silently changes every aggregate underneath them. That is not carelessness on the left — it is the most careful thing available without a comparator. A useful control plane must be able to say no. SparkPilot's approval means something because it can withhold approval.

03 / PROOF

We seeded a rounding difference to prove the comparator can fail.

Every figure in this record is recomputable from the evidence trace. The negative control is deliberate: a comparator that never returns a failure is not a comparator. The clean run of the same workload returns EQUIVALENT, 10 of 10 matching.

EVALUATION / 2026-08-14COMMIT b10aac5
tpch-q1-reference-v1
SCOPE / 1 WORKLOAD RUN ON BOTH PLATFORMS · NO ESTATE INVENTORIED · NO SOURCE SCANNED
BASELINE
emr-6.15.0
Spark 3.4.1-amzn-2
JOB RUN 00g7nppqou31tg0b
CANDIDATE
emr-7.5.0
Spark 3.5.2-amzn-1
JOB RUN 00g7nppt0p4fu80b
DURATION48.5 s126.0 s+159.8%
METERED COST$0.0157$0.0172+9.6%
SUM_CHARGE TYPEdecimal(38,6)decimal(38,1)
EQUALITY
CHECKS
q1_input_sample
matchedschema_typeMATCH
matchedrow_hashMATCH
matchedrow_countMATCH
matchednull_countsMATCH
matchedaggregate_valueMATCH
EQUALITY
CHECKS
q1_pricing_summary
×differedschema_type1 finding · sum_charge: decimal(38,6) → decimal(38,1)
×differedrow_hash1 finding · 984eb4735fc1b990… → 5e7671c91ac6d5cc…
matchedrow_countMATCH
matchednull_countsMATCH
×differed
aggregate_value3 findings
sum_charge:max 1418207490.410181 → 1418207475.3
sum_charge:min 1391007086.081212 → 1391007105.8
sum_charge:sum 8427644657.016188 → 8427644818.5
ASSERTIONS
10
CHECKS DIFFERED
3
FINDINGS
5
JOBS INVENTORIED
0
STATED CAPACITY
UNKNOWN
VERDICT
NOT_EQUIVALENT
DO_NOT_CUT_OVER
MEASUREMENT QUALITY / WALL CLOCK NOT ATTRIBUTABLE TO ENGINE
SINGLE OBSERVATION · NOT BENCHMARK CONTROLLED

04 / THE OPERATING LOOP

Assurance isn't a separate tool. It's the last stage of the loop.

The same control plane that dispatches the run is the one that measures it, compares it, and decides whether it may be promoted.

  1. 01
    CONNECT
    Cross-account role, scoped.
  2. 02
    RUN
    Gates, then dispatch.
  3. 03
    OBSERVE
    State, duration, cost state.
  4. 04
    DIAGNOSE
    Cause, evidence, remediation.
  5. 05
    CHANGE
    Baseline and candidate, paired.
  6. 06
    VERIFY
    Differences are named.
  7. 07
    PROMOTE / BLOCK
    A verdict, with provenance.

05 / THE PRODUCT

Twelve surfaces. One control plane, one audit path.

Runtime-change assurance is the sharpest proof point, not the product. Everything below submits through the same preflight, the same budget gate, and the same audit trail.

NOTHING ASSERTS BEFORE ITS EVIDENCE · THE VERDICT LANDS LAST

06 / GOVERNANCE

The cheapest run is the one that never dispatched.

Policies, budgets, and scopes are evaluated at preflight — before AWS spends a cent. A hard block rejects the submission; a soft warning still lands in the audit trail.

POLICIES
PREFLIGHT / etl-nightly
BLOCKED
ALLOWED INSTANCE TYPESPASS
ALLOWED EMR RELEASE LABELSPASS
MAX VCPU PER RUNHARD BLOCK
REQUIRED SPARK TAGSSOFT WARN
SCOPE / GLOBAL → WORKSPACE → ENVIRONMENT
AWS COST INCURRED / $0
BUDGET GUARDRAILS
ENVIRONMENT / analytics-prod
$8,412 / $10,000 MO
WARN AT 80%CROSSED
BLOCK AT 100%ARMED
SPEND STATE / ESTIMATED + RECONCILED CUR
ACCESS
TEAM ↔ ENVIRONMENT SCOPES
data-platformprod · staging
analyticsstaging
ml-researchsandbox
contractor-extsandbox · user role
IDENTITY / M2M OAUTH · WORKSPACE-SCOPED
EVERY DECISION / WRITTEN TO AUDIT TRAIL
ILLUSTRATIVE NAMES AND AMOUNTS · THE RULE TYPES, THRESHOLDS, ROLES AND SCOPES ARE THE PRODUCT'S

07 / WHERE RESULTS LAND

A verdict nobody sees isn't governance.

Subscribe a channel to run and system events. Delivery is admin-configured and the webhook is stored encrypted.

run.failed · run.timed_out · run.dispatch_failed
run.preflight_failed · run.schedule_submission_failed
evaluation.evaluated · evaluation.revalidation_required
activation.refused
SLACK
AVAILABLE
#data-platform-alerts
NOT_EQUIVALENTtpch-q1-reference-v1, emr-6.15.0emr-7.5.0
3 of 10 checks differed · 5 findings · evidence trace attached
ADMIN-ONLY · WEBHOOK STORED ENCRYPTED · EXAMPLE MESSAGE
EMAIL
OFF BY DEFAULT
Delivered the same way, through Resend, once an operator enables it for the deployment. Listed with its switch, so a default is never mistaken for a feature.

08 / THE ENVIRONMENT

We can build it. We can delete it. Or you can bring your own.

Three engines, three genuinely different shapes of commitment. The control stays with your engineers in all three.

Whichever you choose, the control plane only ever assumes a role into your account. Compute, data, and evidence never leave it.

PATH 01
EMR Serverless
NOTHING TO PROVISION
The fastest path from signup to a first run — because there is nothing to build.
No cluster, no node pools, no capacity to keep warm, nothing billing while idle. And no teardown, because nothing is standing.
PROVISIONING
NONE · YOU BRING THE APPLICATION
IDLE COST
$0 BY CONSTRUCTION
AUTO-STOP
15 MIN ON APPS WE CREATE
TEARDOWN
NOTHING TO TEAR DOWN
PATH 02
EMR on EKS
WE BUILD IT, OR YOU BRING IT
One click builds the whole environment in five resumable stages — and one action destroys it.
Or point us at the EKS cluster you already run, and we use it as it is. Cheapest at steady state with spot pricing.
IF WE BUILD IT · FIVE RESUMABLE STAGES
01
NETWORK
VPC, two AZs, a NAT gateway per AZ.
02
EKS
1.35 in Auto Mode, IRSA and Pod Identity.
03
EMR
Virtual cluster with a scoped execution role.
04
BOOTSTRAP
Validate what was built before using it.
05
RUNTIME
A real job runs before it is handed over.
CAPACITY
DRIVERSON-DEMAND
EXECUTORSSPOT-FIRST
Two Karpenter node pools. No static node groups — Karpenter owns capacity, so there is nothing to keep warm.
TEARDOWN
terraform destroy
Every resource the build created, removed in reverse. No orphaned NAT gateway quietly billing for a year.
PATH 03
EMR on EC2
LONG-LIVED, AND YOURS TO SHAPE
For teams that want the cluster to persist and want to own instance selection.
We detect when no auto-termination policy is set and offer to set the native AWS one — rather than building a scheduler of our own that drifts out of step with it.
CLUSTER LIFETIME
YOURS TO DECIDE
IDLE COST BETWEEN CLUSTERS
$0 — NOTHING STANDS
AUTO-TERMINATION
DETECTED · NATIVE AWS POLICY

And in every path, the boundary is the same.

SparkPilot dispatches and measures. The compute, the data, and the evidence stay in the account they already live in.

PROVISIONING / TERRAFORM · RESUMABLE
ACCESS / ASSUMEROLE · EXTERNALID
AUTH / M2M OAUTH

Engines today: EMR on EKS, EMR Serverless, EMR on EC2.

09 / COST

Money is decided at four moments. Three of them happen before the run.

A dashboard shows you what you already spent. Most of the leverage is upstream of it — in the engine you pick, the shape of the create call, and the hours nobody is watching.

MOMENT 01
BEFORE YOU COMMIT

Engine choice, priced — while the choice is still free.

At environment creation, before a dollar is committed, the product prints the monthly floor of each engine with the arithmetic attached. For a low-volume team this single screen is the largest lever in the product.

EMR on EKS
~$152/mo
EKS control plane $73.00
+ smallest practical always-on capacity
8 runs/month ≈ $152.10/mo
EMR Serverless
$0/mo
No standing compute.
You pay for the seconds a job runs.
8 runs/month ≈ $0.104/mo
EMR on EC2
$0/mo
Between clusters, nothing stands.
Cost begins when a cluster does.
Per-cluster hours, priced on create
MOMENT 02
AT CREATION

Cost-correct by construction.

Cost correctness is a property of the create call, not something the customer has to remember afterwards. This is how SparkPilot's own EMR Serverless create is written; the product today attaches to an application you bring.

·initialCapacity omitted — zero idle cost by construction, not by policy.
·Auto-stop at 15 minutes, set at birth.
·A maximumCapacity hard ceiling, so a runaway job has a bound.
·No VPC attachment unless asked for — which avoids per-GB NAT charges entirely.
·Cost-allocation tags applied at birth, so CUR lines are labelled retroactively rather than orphaned.
MOMENT 03
WHILE IT RUNS

Cheaper by default, and it refuses the unbounded case.

·Spot-first executors with on-demand drivers and graceful decommissioning — the executor fleet at Spot rates, not on-demand.
·Graviton / arm64 preferred, with the 20% discount applied to every estimate rather than promised in a footnote.
·Karpenter consolidates empty nodes within a minute of the last task.
·A rate-card cost ceiling is shown and confirmed before submit — and the submission is refused outright when dynamic allocation is unbounded.
·In-flight budget reservations, so two concurrent submissions cannot each spend the budget the other is already using.
·Quotas and a hard queue-capacity gate ahead of dispatch.
MOMENT 04
WHILE IT SITS IDLE

The money nobody watches.

We measured a real idle Spark environment for a month. It cost $230.53. Five cents of that was Spark.

ONE CLUSTER · NOTHING SCHEDULED · OUR REFERENCE ENVIRONMENT · JULY 2026
IDLE ENVIRONMENT · 30 DAYSus-east-1
Unused node capacity$94.45
EKS control plane — no idle action reduces it$85.34
Control-plane log ingestion, never expiring$22.61
Everything else metered on the account$28.13
TOTAL BILLED$230.53
OF WHICH, ACTUAL SPARK WORK$0.05

This bill does not get smaller when you run less. It is what the cluster costs to exist, and you pay it once per cluster — dev, staging, prod, per team, per region. At 30 clusters, the per-cluster lines above come to $6,072 a month of standing charge before a single job is submitted — an illustration of how a per-cluster charge scales, not a measurement of anyone's estate. The account-level remainder is excluded, because it does not recur once per cluster.

THE SAME CLUSTER, TIMES YOUR ESTATEestimated, not measured
Disable the audit control-plane log type, and set retention — needs UpdateClusterConfig, which SparkPilot does not request$20.61
System node pool from two nodes to one — needs no high-availability requirement on this cluster$48.77
NAT gateway to an S3 gateway endpoint — needs nothing else needs the NAT path$7.35
Floor no lever reaches (the control plane)$85.34
ESTIMATED AVOIDABLE PER CLUSTER, PER YEAR$921
AT 30 CLUSTERS, PER YEAR$27,623

Two thirds of that is the node lever, and it is only available if this cluster does not need high availability — one system node is a single point of failure for cluster DNS. A cluster that must stay highly available avoids $27.96 a month, not $76.73.

Pull every lever and this cluster still bills $153.80 a month while idle. SparkPilot surfaces these three with the command to run; you execute them in your own account, because performing them needs permissions we do not ask for — disabling audit logging requires a permission that would also let us expose your Kubernetes API server. Retention alone does not remove the log charge: it is ingestion, not storage.

Read your own number rather than ours: it is split_line_item_unused_cost in your Cost and Usage Report. That is the field this page is built on, and it is the field that would prove us wrong.

THE CONTROLS
Stop, restore, and schedule idle capacity — each behind an admin capability, declined by default.
A customer-installed stack that defaults to ReportOnly and only ever acts on clusters you have tagged.
Notebook sessions stop themselves at 30 idle minutes, with the avoided spend recorded.
WHAT NO CONTROL CAN FIX
The EKS control plane bills whether or not a job ever runs. Deleting the environment is the only lever — which is why deletion is a first-class action, not a support ticket.
90 DAYS AHEAD OF THE CLIFF

When your EKS version leaves standard support, the cluster control-plane fee goes up six times.

Extended support adds $0.50 per hour to the $0.10 cluster fee: $0.10 to $0.60 per cluster-hour, charged whether or not anything is running. Your Spark compute is billed exactly as before — this is the control-plane fee, not the workload. We warn you 90 days out, with the exact date.

EKS 1.34 · STANDARD SUPPORT ENDS
2026-12-02
TODAY$0.10/hr · $73/mo
AFTER$0.60/hr · $438/mo
Control-plane fee only, per cluster, us-east-1 list price at 730 hours/month. Dates are UTC, per the AWS release calendar. Version dates and rates read from the AWS EKS and Pricing APIs on 2026-09-18.
AND WHAT WE WON'T DO
NO FABRICATED $0.00
An unknown renders unavailable. A cost that reads $0.00 because we weren't allowed to look is indistinguishable from one that reads $0.00 because nothing was spent.
GROSS, NOT NET
AWS credits covered all of it, so you were charged $0.00 — but credits are finite, and this is the figure you'll pay when they run out.
WE DON'T REIMPLEMENT AWS
On EMR on EC2 we detect that no auto-termination policy is set and offer to set the native one — rather than building a scheduler of our own that drifts out of step with it.
AND EVERY RUN'S COST MOVES THROUGH THREE STATES
ESTIMATED
~$47.80
PRICING MODEL · PRE-DISPATCH
PENDING
NOT MEASURED
AWAITING CUR · NEVER $0.00
RECONCILED
$47.82
CUR · 03:12 UTC
ILLUSTRATIVE AMOUNTS · THE THREE STATES ARE HOW EVERY RUN IS LABELLED

10 / INTEGRATES INTO YOUR STACK

SparkPilot does not replace your orchestrator.

Your scheduler keeps scheduling. The gates, the cost state, and the evidence come from here. Four interfaces into the same control plane.

01
Install the provider on your scheduler — from source during a pilot; the package is not on PyPI yet.
pip install apache-airflow-providers-sparkpilot # preview
02
Add a SparkPilot connection — workspace URL and M2M client credentials.
airflow connections add sparkpilot_default …
03
Swap the operator. Preflight, budget gate, and evidence come with it.
SparkPilotSubmitRunOperator(job_id=…)
INSTALL BOUNDARY
Runs on your Airflow scheduler. Nothing is installed in your Spark cluster, and DAG definitions stay in your repository. Preview: exercised against test stubs; its first real scheduler-driven run is on the validation list.
10 / WHAT IS NEXTROADMAP · NOT SHIPPED

Run:ai gates chips. Kueue gates cores. Nobody gates dollars.

Everything below is a destination, not a feature. It is on this page because a roadmap you can read is worth more than one you have to ask for — and because we would rather you hold us to it.

GPU AND INFERENCE WORKLOADS
ROADMAP

The same verdict, dispatch and reconcile path applied to training and inference spend. GPU quota tools today are denominated in chips, not dollars, and none of them reconcile against the bill. That gap is the reason this is first on the list.

RAY / KUBERAY
ROADMAP
Distributed Python beside Spark, under the same policy and budget gates.
SAGEMAKER
ROADMAP
Job governance for training and processing runs.
BEDROCK
ROADMAP
Token and inference spend under the same budget substrate.
DATABRICKS
ROADMAP
Jobs API routing, so a governed run does not care which platform executes it.
SNOWFLAKE
ROADMAP
Warehouse spend inside the same reconciliation.
VERDICT CLI / GITHUB ACTION
ROADMAP
The gate at pull-request time, before the DAG ever merges.
AGENT-SAFETY API
ROADMAP
The same pre-execution gate, exposed to agents rather than people.

11 / PRICING

Four ways in. The gaps are printed.

A dash means the capability is not in that tier. We would rather lose the deal than discover it during the pilot.

CAPABILITY
Scan
Free
Assessment
$12,500
Assurance
from $24,000/yr
Platform
from $750/mo
Portability review of job definitions and source✓ Included✓ Included✓ Included✓ Included
Constructs with no target equivalent, and what they cost to fix✓ Included✓ Included✓ Included✓ Included
Execution on both source and target platformNot included✓ Included✓ IncludedNot included
Row-level output comparisonNot included✓ Included✓ IncludedNot included
Measured wall clock and resource consumption, both sidesNot included✓ Included✓ IncludedNot included
Cutover recommendation with coverage countsNot included✓ Included✓ IncludedNot included
Re-checked on every runtime or config changeNot includedNot included✓ IncludedEnterprise rung
Governed dispatch in your own AWS accountNot includedNot includedNot included✓ Included
Preflight, policy and quota gates before dispatchNot includedNot includedNot included✓ Included
Per-run audit trail and failure diagnosticsNot includedNot includedNot included✓ Included
Scan
Free
  • Portability review of job definitions and source
  • Constructs with no target equivalent, and what they cost to fix
  • Execution on both source and target platform — not included
  • Row-level output comparison — not included
  • Measured wall clock and resource consumption, both sides — not included
  • Cutover recommendation with coverage counts — not included
  • Re-checked on every runtime or config change — not included
  • Governed dispatch in your own AWS account — not included
  • Preflight, policy and quota gates before dispatch — not included
  • Per-run audit trail and failure diagnostics — not included
Assessment
$12,500
  • Portability review of job definitions and source
  • Constructs with no target equivalent, and what they cost to fix
  • Execution on both source and target platform
  • Row-level output comparison
  • Measured wall clock and resource consumption, both sides
  • Cutover recommendation with coverage counts
  • Re-checked on every runtime or config change — not included
  • Governed dispatch in your own AWS account — not included
  • Preflight, policy and quota gates before dispatch — not included
  • Per-run audit trail and failure diagnostics — not included
Assurance
from $24,000/yr
  • Portability review of job definitions and source
  • Constructs with no target equivalent, and what they cost to fix
  • Execution on both source and target platform
  • Row-level output comparison
  • Measured wall clock and resource consumption, both sides
  • Cutover recommendation with coverage counts
  • Re-checked on every runtime or config change
  • Governed dispatch in your own AWS account — not included
  • Preflight, policy and quota gates before dispatch — not included
  • Per-run audit trail and failure diagnostics — not included
Platform
from $750/mo
  • Portability review of job definitions and source
  • Constructs with no target equivalent, and what they cost to fix
  • Execution on both source and target platform — not included
  • Row-level output comparison — not included
  • Measured wall clock and resource consumption, both sides — not included
  • Cutover recommendation with coverage counts — not included
  • Re-checked on every runtime or config change — Enterprise rung
  • Governed dispatch in your own AWS account
  • Preflight, policy and quota gates before dispatch
  • Per-run audit trail and failure diagnostics
AWS INFRASTRUCTURE IS BILLED TO YOUR ACCOUNT, NOT RESOLD BY US · THE FULL PRICING PAGE

12 / GETTING STARTED

Bring the change you don't trust.

SPARKPILOTSPARK CHANGE WITH EVIDENCE