Every labeling project runs into the same wall eventually: two annotators look at the same piece of data and make different calls. That gap is what annotation guidelines exist to close. They spell out how edge cases get handled, what counts as ambiguous, and when to escalate instead of guessing. The strongest guidelines get tested against a small batch before the full dataset goes through, since that's when gaps between what's written and how it's actually applied tend to show up.