Honest Limitations

What SparkPilot does not do yet

Our product gates claims about your jobs, so we gate claims about our product. This page is derived from the same internal claims register that backs a CI tripwire on every page of this site, with human review behind it. If something here contradicts our marketing copy, the marketing copy is wrong — tell us.

Migration and assurance scope

What our migration and equivalence evidence actually covers. If you arrived from the pricing page, this is the fine print behind it.

One workload, one runtime-upgrade pair — that is the whole proof

The only migration we have executed end to end is an EMR Serverless runtime upgrade, emr-6.15.0 to emr-7.5.0, on a single reference workload in a real AWS account. It compared clean across 10 equality assertions. That is the entire body of executed evidence behind everything the pricing page says about moving between platforms.

Jobs inventoried: 0. Source files scanned: 0.

Those two lines are copied verbatim from the coverage block of that same report, and they are the honest state of the estate inventory and the static scanner: neither has ever been run against a customer estate, so neither is proven. Early compatibility scans are read by an engineer rather than produced by a tool. If we ever show you a scan, it will print its own coverage counts so you can see what it did not look at.

There is no catalogue of supported migration paths

One source-and-target pair has been executed. Any other pair — a different engine, a different vendor, a different runtime jump — is unproven until we run it. We will not publish a support matrix for migrations we have not performed, and no page on this site should imply one exists.

The verdict is deliberately scoped, and it can say no

The published reference report's recommendation is INSUFFICIENT_EVIDENCE: the one assessed pair was equivalent, but no estate inventory or static scan was attached, so the report refuses to recommend cut-over. PROCEED_FOR_ASSESSED_SCOPE is issued only when inventory, scan, and dual run all pass, and even then it says nothing about workloads that were not run. It is also not a rubber stamp: a negative control run with a deliberately seeded rounding difference returned DO_NOT_CUT_OVER, which is what stops 'we found no differences' from being a comparator that always agrees. Continuous assurance re-runs this same mechanism and inherits every limit on this page.

Current scope limits

These are real boundaries of the product today, stated so your evaluation does not discover them for us.

An idle environment still bills, and deleting it is the only lever

Your EKS control plane bills about $0.10/hour — roughly $2.40 a day, $73 a month — from the moment the cluster exists until it is deleted, whether or not a single Spark job runs. Worker nodes, EBS, NAT and data transfer add to that when work is running. SparkPilot shows this idle rate on the environment so it is not a surprise, but we do not yet offer stop, suspend or scale-to-zero: today the only way to stop paying for an idle environment is to delete it. Per-run figures in the product price the Spark work only, and are typically cents against these infrastructure costs.

IAM simulation runs on EMR on EKS only

Pre-dispatch iam:SimulatePrincipalPolicy checks cover EMR on EKS with BYOC-Lite onboarding. On EMR Serverless and EMR on EC2, preflight enforces policy and quota gates but does not simulate per-action IAM permissions before submit.

Engine coverage is uneven beyond EMR on EKS

EMR on EKS is the end-to-end governed path today, with completed runs re-verified against AWS. EMR Serverless and EMR on EC2 can be created now; on each, the engine-level provision, submit and in-flight cancel path passed once, on 2026-06-01, before the current dispatch architecture. What has not been proven on those two engines is governed dispatch through the current control plane and steady-state reconciliation under it.

Budget blocking works; its failure explanation does not yet

Per-team monthly budgets warn and hard-block dispatch — an over-budget run was blocked before dispatch in a live account, with zero AWS spend. The rough edge: the blocked run's in-product explanation currently misattributes the cause to environment setup instead of the budget, and its diagnostics panel is empty. The block is real; the explanation is being fixed.

CUR reconciliation is early access

Billed-truth reconciliation against your Cost and Usage Report — down to Split Cost Allocation Data on EKS — is built and has been exercised end to end against a simulated report; no billed run has yet been reconciled through a real CUR, and the production onboarding path is not yet reliable. It ships as early access until both are proven in production.

Per-run cost attribution on EMR Serverless is not CUR-proven

Cost-allocation tag propagation for Serverless is built, but the tag-to-CUR proof so far covers EMR on EKS only. Per-run billed attribution on Serverless is coming soon. On EMR on EC2, per-run billed attribution is not possible at all: EMR bills the cluster, and a step is not a billable resource.

Airflow and Dagster providers are preview

Both providers are installable from source and exercised against stubs in tests. A first real scheduler-driven run through a customer-style deployment is still on the validation list — we do not call them GA.

No dead-letter handling for failed dispatch

The run lifecycle is idempotent and transactional, and failures are classified with diagnostics. But there is no dead-letter queue replay: after certain infrastructure failures a run must be re-submitted. We do not claim that no run is ever lost.

The AI assistant is not shipped

Run diagnostics and exact-fix remediation are rule-based today. Nothing on this site should read as a shipped AI operator, and if it does, report it to us.

Not supported today (roadmap only)

We do not run these workloads or integrate with these platforms today. They appear on our expansion roadmap, and nowhere on this site should they be described as supported.

GPU and inference workloads

No GPU scheduling or inference serving today; expansion-map item.

Ray / KubeRay

Spark engines only; no Ray cluster support.

SageMaker

No SageMaker job governance.

Bedrock

No Bedrock usage governance.

Databricks

Jobs API routing is planned on our roadmap; nothing ships today.

Snowflake

No Snowflake integration.

Multi-cloud

SparkPilot is AWS-only by design for now.

Evaluating us anyway?

If the limits above are acceptable for your workload shape, the parts that do ship are proven with live evidence — IAM preflight, policy and quota hard-blocks, governed dispatch on EMR on EKS, and per-run audit. Bring a workload and check.