Description #

The methodology paper behind provider level claims anomaly detection: a supervised learning study comparing Random Forest, Support Vector Machine, and K Nearest Neighbors classifiers on structured Medicare style claims data, at a fraud prevalence of about 9 percent. The best configuration, a Random Forest with a depth of 25 and 1,000 estimators, reached 86.88 percent accuracy, an F1 score of 0.819, and an AUC ROC of 0.94. Covers the feature engineering behind that result, including chronic condition indexing and derived survival status, and examines the regulatory framework a detector operates inside: the False Claims Act, the Anti Kickback Statute, the Stark Law, and HIPAA.

Files #

PathSizeSHA-256
files/healthcare-provider-fraud-detection.pdf 280,469 bytes b10524c58e4e6ed75531f9f777e84b50357603433aaaeaffcdcb2cf3bf634517

Each file's SHA-256 is listed above. To confirm a download is unmodified: shasum -a 256 filename

Licence: CC BY 4.0

← Back to Initiative 1