kubeflow/sdk
View on GitHubSparkConnect CR cleanup behavior on failure and SDK state mismatch
Open
#476 opened on Apr 30, 2026
area/sparkgood first issuehelp wantedkind/bug
Repository metrics
- Stars
- (124 stars)
- PR merge metrics
- (PR metrics pending)
Description
Summary
While working on port-forward cleanup in the SDK #460 , it surfaced that there may be a broader gap in how SparkConnect resources are handled on failure.
Observations
- SparkConnect CR does not appear to clean up resources when connection/setup fails
- This can leave resources running if failure happens mid-connection
There is a related issue in the Spark Operator: kubeflow/spark-operator#2927
Additional Observation
There also appears to be a mismatch in state mapping:
- The SDK exposes a "Running" state
- SparkConnect CR does not expose a corresponding state