Thanks for the feedback!
My total goal isn’t just for a single metric to decide whether an episode should be deleted. Instead, I want Calibra to surface episodes that are worth investigating before training whether because they indicate a recording issue or because they are genuinely interesting edge cases.
The workflow you described (compute diagnostics > inspect a shortlist > decide what to do) aligns well with the direction i am aiming . I will take a look at LeRobot issue #3760.