I. Introduction
Imagine you are visited by police officers, not because of anything you did, but because an algorithm decided you were likely to. This is not a story out of a dystopian fiction novel but an actual incident which took place in Chicago in 2013, when the city’s police department used a predictive ‘heat list’ to flag hundreds of young Black men as probable perpetrators of gun violence, many with no criminal history at all.
Predictive policing uses AI and data analytics to forecast criminal activity before it occurs. Drawing on vast datasets of arrest records, criminal histories, geographic information, social media activity, and biometric surveillance feeds, these systems build probabilistic models of future behaviour and direct police resources accordingly. According to a recent report, smart technologies could help cities reduce crime by 30 to 40 per cent and cut emergency response times by 20 to 35 per cent. Yet this expansion consistently violates ethical and legal safeguards, creating ambiguity and a dangerous gap between what these tools promise and whom they actually harm. At the centre of this gap lie two interconnected crises: algorithmic bias and the opacity of such algorithms. This article argues that these two are dimensions of the same structural injustice, and that structural reform, not outright abolition, is the way forward to harness the benefits of algorithmic policing without letting it become a tool for discrimination.
II. The Problem and Concerns: Dirty Data, Biased Algorithms, and the Black Box
Societal discrimination is as significant a problem as algorithmic bias itself. According to a study from the New York University School of Law (NYU Law Review), the majority of the predictive policing data sets studied by the researchers during a particular case study suffer from “dirty data” — data that has been influenced not just by historical racism but also by a “history of police scandals” and deliberate manipulation of official crime records, collectively characterised by the study as‘ juking of stats’. An algorithm that uses such data will learn to perpetuate and magnify the current level of bias by identifying the same communities as high-risk and sending additional law enforcement to those communities which, in turn, generates additional data that confirms their risks. This creates a self-sustaining cycle of discrimination while appearing to be based upon objective data.
Several independent think tank studies show that Black people are five times more likely to be arrested than white people, a disparity that, when encoded into training data, becomes the algorithm’s baseline. This is not a flaw that can be fixed by improving the algorithm’s mathematics, but a flaw inherited from the society’s racial biases that produced the data in the first place.
Adding to this structural flaw is opacity. Most predictive policing tools are built by private companies that guard their algorithms as trade secrets, the so-called ‘black box‘ problem. Citizens subject to algorithmic targeting, therefore, have no meaningful avenue to contest decisions that the rule of law would ordinarily require to be justified, since the whole process takes place without ever revealing how, or from what data, the prediction was generated, and this is not uniquely a Western pathology.
Delhi’s CMAPS (Crime Mapping Analytics and Predictive System) collects data every 3 minutes from the Indian Space Research Organisation’s satellites, historical crime data, and the ‘Dial 100’ helpline, to identify ‘crime hotspots’ which are crime-prone areas. Hyderabad police go a step further as they use sensitive data taken from the ‘Integrated People Information Hub’ – containing family details, biometric details, passport details, addresses, and even bank transaction details – to determine who is more likely to commit crimes. However, both these organisations pose similar challenges, as neither the algorithm used, nor the importance given to each data source has ever been published, making it impossible to verify whether these tools are actually identifying potential crime, or simply validating the old data, just like the feedback loop the NYU study documented in the 13 US jurisdictions.
The potential issue that minorities might be targeted with the help of AI systems is not merely a hypothetical, but an unfortunate reality. In March 2020 in Delhi, when arrests connected to the infamous Anti-CAA riots were being made, a young man named Ali (name changed) had his life turned upside down when AI–powered facial recognition technology was used to identify him as a culprit in the Ratan Lal Murder Case (FIR 60/2020). The investigating officer told media agencies that the case was “solved” using image enhancement tools (Amped FIVE by Amped Software) and facial recognition software (AI Vision by Innefu Labs). This however is the same case where 27 of the 29 accused have been granted bail, and one other has been discharged. . According to the defence counsel, the methodology of facial recognition was not disclosed before the court, nor were the claims of the cases being “solved” by such technologies backed by verifiable proof. In fact, the same FRT (Face Recognition Technology) infrastructure had, in an unrelated 2018–19 Delhi High Court proceeding, been disclosed as returning matches at accuracy rates of merely 1–2 percent. Then comes the second structural flaw of predictive policing systems, i.e. dirty data. Research by the Vidhi Centre for Legal Policy has found Delhi’s CCTV and police-station density is itself structurally skewed toward over-policed, Muslim-majority neighbourhoods, meaning the footage feeding these systems was compromised before any algorithm touched it — and because nothing about the match was ever disclosed, no one had a way to test whether that compromise had produced a wrongful identification until it was too late.
III. Structural Solutions: Reform, Not Abolition
The problems of predictive policing are structural and demand structural solutions. The root cause is dirty data rather than a fatal flaw in the concept itself, meaning reform, not abolition, is the most viable path, and that such a path is necessary and achievable. This means that the data sets must not just clean themselves of the biased data, but must also be absolutely transparent about the whole process.
The first required reform is transparent data collection subject to independent public oversight. Existing datasets must be audited for discriminatory patterns, weighted to correct for historical over-policing, and regularly updated rather than treated as static historical records. Alongside this, ‘white-box’ algorithms must be mandated, i.e., interpretable, auditable systems open to independent review by bodies comprising civil rights advocates, data scientists, and community representatives with genuine authority, not merely advisory roles, independent of both law enforcement and private vendors – a standard India’s current framework is unable to meet.
In India, where law enforcement algorithms may quietly be exempted from Data Protection Board’s audits through a mere executive notification by virtue of Section 17(2)(a) of the DPDPA, what is required therefore is a dedicated algorithmic oversight body established by separate legislation — analogous in institutional design to the National Human Rights Commission under the Protection of Human Rights Act, 1993, whose mandate explicitly covers law enforcement data systems and whose independence from both the executive and private vendors is constitutionally secured. Such a body may consist of a mix of retired High Court or Supreme Court judges, a technical expert in algorithmic auditing, a civil society representative, and a data protection officer. However, NHRC’s biggest weakness – its reliance on state compliance, can serve as a warning for the body. Instead of being relegated to a recommendation committee, the algorithmic oversight body should be trusted with statutory powers which may include a mandatory pre-deployment audit of any predictive system and power to regulate private vendors involved with such technologies.
The ICO in the UK offers a useful institutional model, acting as a statutory body with genuine investigative authority over surveillance and privacy concerns, rather than holding merely an advisory role.
Additionally, the legal frameworks governing challenge and remedy must be strengthened. Individuals targeted by predictive systems must have a right of access to the factual basis of the algorithmic decision that affected them, analogous to the right of an accused to confront the evidence against them. Courts should develop a doctrine that treats discriminatory algorithmic outputs as triggering constitutional equal protection scrutiny regardless of intent, recognising that the equal protection guarantee applies to systems as well as individual officers.
Finally, and most fundamentally, workable solutions must centre the communities most harmed. Community-centred design must start with building systems with, not merely for, those historically over-policed. Trust between law enforcement and citizens is a prerequisite for effective policing, not an optional extra. Systems that surveil without consent, predict without transparency, and act without accountability, corrode precisely the social foundation on which policing depends.
IV. Conclusion
We, as a society, are standing at crossroads. Predictive policing promises us safer cities, smarter law enforcement, and data-driven justice. It is a compelling promise. But promises made by machines are only as good as the humans who built them, and the data they were built on.
The convergence of data protection failures and algorithmic bias is not accidental. For generations, certain communities have been watched more closely, questioned more frequently, arrested more often, not because they commit more crime, but because policing has more often than not followed power, and power has rarely protected the powerless. Those arrest records are what we are feeding into these algorithms, and then we are surprised when the machine tells us to watch the same people, in the same neighbourhoods, all over again. This is not a glitch, but a system working exactly as designed. The oversight body proposed in the potential reforms is not just an abstraction, for had it existed in 2020, its mandatory pre-deployment audit would have required Delhi Police to disclose the FRT’s accuracy and matching basis before Ali’s identification could stand as evidence at all, and he is not the only one who has faced gross injustice at the hands of an erroneous algorithm.
But none of this is inevitable. Transparent data, accountable algorithms, independent oversight, and real democratic accountability to the communities most affected are not radical demands — they are the minimum any just society should require before granting a piece of software the power to mark a person as a future criminal.
Kunj Bhatia is a second-year law student at NLU Jodhpur. His areas of interest include the intersection of technology and constitutional rights, alongside their political and socio-economic implications.

