“Which dataset actually trained this model?”
You trained on a dataset a vendor sent you in March. In May, someone edited three rows and re-sent the “same” file.
Nothing in your pipeline noticed. The checksum you kept was for the old version — and even if it had failed, it would only have told you something changed. Not what. Not when. Not who.
Can you prove which file trained your model to an auditor who has no reason to trust your infrastructure?