Who’s bogus? Learning to identify fraudulent firms from unbalanced and one-side labelled tax return data

Journal article COMPASS '18: Proceedings of the 1st ACM SIGCAS Conference on Computing and Sustainable Societies Tax and Tax compliance

Abstract

We apply a classifier to the value added tax (VAT) returns from Delhi (India) to increase tax compliance by identifying "bogus" (shell) firms which can be further targeted for physical inspections. We face a nonstandard applied machine learning scenario. First, onesided labels: firms that are not caught as bogus are of unknown class: bogus or legitimate, and we need to not only use them to train the classifier but also make predictions on them. Second, multiple time-periods: each firm files several periodic VAT returns but its class is fixed so prediction needs to be made at the firm, not firm-period, level. Third, point in time simulation: we estimate the revenue saving potential of our model by simulating the implementation of our system in the past. We do this by rolling back the data to the state of knowledge at a specific time and calculating the revenue impact of acting on our model's recommendations and catching the bogus firms, and estimate US$40 million in recovered revenue.