Loading...
When labeled examples are scarce, synthetic samples can bootstrap classifiers — but bias and leakage risks remain.
The right use is balancing rare classes, not inventing an entire dataset from scratch. Final validation always belongs on a held-out real set.
If the synthetic distribution drifts from reality, the model looks strong in the lab and weak in the field. Treat distribution-gap reporting as part of delivery.
For customer data, synthesis is not a privacy loophole; document origin and use limits.
Manager action: Use synthetic data only to balance rare classes; always validate on a held-out real set.
