Honest Limitations

What SparkPilot does not do yet

Our product gates claims about your jobs, so we gate claims about our product. This page is derived from the same internal claims register that backs a CI tripwire on every page of this site, with human review behind it. If something here contradicts our marketing copy, the marketing copy is wrong — tell us.

Current scope limits

These are real boundaries of the product today, stated so your evaluation does not discover them for us.

An idle environment still bills, and deleting it is the only lever

Your EKS control plane bills about $0.10/hour — roughly $2.40 a day, $73 a month — from the moment the cluster exists until it is deleted, whether or not a single Spark job runs. Worker nodes, EBS, NAT and data transfer add to that when work is running. SparkPilot shows this idle rate on the environment so it is not a surprise, but we do not yet offer stop, suspend or scale-to-zero: today the only way to stop paying for an idle environment is to delete it. Per-run figures in the product price the Spark work only, and are typically cents against these infrastructure costs.

IAM simulation runs on EMR on EKS only

Pre-dispatch iam:SimulatePrincipalPolicy checks cover EMR on EKS with BYOC-Lite onboarding. On EMR Serverless and EMR on EC2, preflight enforces policy and quota gates but does not simulate per-action IAM permissions before submit.

Engine coverage is uneven beyond EMR on EKS

EMR on EKS is the end-to-end path today. EMR Serverless and EMR on EC2 have dispatch clients, but state reconciliation and cancel propagation do not route to those engines yet.

Budget blocking works; its failure explanation does not yet

Per-team monthly budgets warn and hard-block dispatch — an over-budget run was blocked before dispatch in a live account, with zero AWS spend. The rough edge: the blocked run's in-product explanation currently misattributes the cause to environment setup instead of the budget, and its diagnostics panel is empty. The block is real; the explanation is being fixed.

CUR reconciliation is early access

Billed-truth reconciliation against your Cost and Usage Report — down to Split Cost Allocation Data on EKS — has been proven end to end, but the production onboarding path is not yet reliable and is actively being fixed. It ships as early access until that work is proven in production.

Per-run cost attribution on EMR Serverless is not CUR-proven

Cost-allocation tag propagation for Serverless is built, but the tag-to-CUR proof so far covers EMR on EKS only. Per-run billed attribution on Serverless is coming soon.

Airflow and Dagster providers are preview

Both providers are installable from source and exercised against stubs in tests. A first real scheduler-driven run through a customer-style deployment is still on the validation list — we do not call them GA.

No dead-letter handling for failed dispatch

The run lifecycle is idempotent and transactional, and failures are classified with diagnostics. But there is no dead-letter queue replay: after certain infrastructure failures a run must be re-submitted. We do not claim that no run is ever lost.

The AI assistant is not shipped

Run diagnostics and exact-fix remediation are rule-based today. Nothing on this site should read as a shipped AI operator, and if it does, report it to us.

Not supported today (roadmap only)

We do not run these workloads or integrate with these platforms today. They appear on our expansion roadmap, and nowhere on this site should they be described as supported.

GPU and inference workloads

No GPU scheduling or inference serving today; expansion-map item.

Ray / KubeRay

Spark engines only; no Ray cluster support.

SageMaker

No SageMaker job governance.

Bedrock

No Bedrock usage governance.

Databricks

Jobs API routing is planned on our roadmap; nothing ships today.

Snowflake

No Snowflake integration.

Multi-cloud

SparkPilot is AWS-only by design for now.

Evaluating us anyway?

If the limits above are acceptable for your workload shape, the parts that do ship are proven with live evidence — IAM preflight, policy and quota hard-blocks, governed dispatch on EMR on EKS, and per-run audit. Bring a workload and check.