HomeAI: Scale model or reducerUncategorizedAI: Scale model or reducer

AI: Scale model or reducer

The AI Revolution and Its Hidden Toll

Since the explosive arrival of generative AI in the public sphere, barely a day goes by without a new platform promising to revolutionize how we work and live. « Homo technologicus, » fascinated by these seemingly infinite possibilities, has eagerly integrated AI into daily life without asking too many questions. While environmentalists warn of energy footprints, trade unionists worry about job displacement, and cynics fear a Terminator-style doomsday, their voices are easily drowned out by the sheer enthusiasm surrounding these new tools. But as artificial intelligence steadily conquers territories once reserved exclusively for human experts—such as analyzing and generating software test cases—it is time to take a step back and examine the profound risks hidden beneath the surface.

The Bias Trap: Why « Fixing » AI is Harder Than It Looks

When we discuss AI risks, bias and ethical failures are usually the first to be highlighted. These flaws stem from narrow data models, poor training, bad labeling, or opaque algorithms. The real-world consequences are already well-documented:

  • Predictive Policing Flaws: Algorithms used for criminal recidivism risk assessments in the US and predictive policing in Chicago have shown clear, systemic bias.
  • Weapons of Math Destruction: As mathematician Cathy O’Neil details in her book, unchecked mathematical models can actively reinforce inequality and damage social systems.
  • Opaque Governance: The controversial departure of ethics researcher Timnit Gebru from Google highlighted the corporate sensitivity and lack of transparency surrounding Large Language Models (LLMs).

Many assume that once a bias is discovered, we can simply patch it. Unfortunately, it is not that simple. For every bias we uncover, dozens remain buried deep within the data models, waiting for a specific prompt to reveal them.

Did you know? Attempting to « force » inclusivity can backfire. When Google’s Gemini (formerly Bard) tried to self-correct for racial bias, it overcompensated, producing bizarre historical anomalies like black Vikings and female Popes. Balancing these models against rapidly evolving societal standards is an incredibly complex task.

Compounding this issue is the « black box » nature of proprietary AI. Tech giants guard their algorithms as closely held trade secrets to prevent fraud and protect intellectual property. However, this total lack of explicability prevents us from truly understanding or learning from past failures, leaving users with a repetitive cycle of « it will work better in the next version. »

The Danger of Undetected Hallucinations and Blind Trust

In risk management, we define risk as the product of probability and impact. However, risk is heavily mitigated if our ability to detect the failure is high. For example, the risk of a car leaving the assembly line with a wheel missing is virtually nil because the defect is glaringly obvious.

Similarly, if an AI generates an image of a person with six fingers, the error is immediately detectable. But what happens when the errors are subtle? If an AI generates software test cases designed to check that a negative price never appears on a train ticket, a non-expert might easily trust the output blindly without realizing the logic is flawed.

The real danger of AI hallucinations is not the hallucination itself—it is our blind trust in the results and the absence of human expert validation.

The Rise of Sophisticated Deception

Beyond accidental errors, we must also contend with deliberate hallucinations. Highly realistic deepfakes and real-time audio/visual manipulation are becoming incredibly difficult to detect. In a hyper-connected, fast-paced world where immediate consumption trumps fact-checking, these technologies present unprecedented challenges for verification and quality assurance.

Model Collapse and the Erosion of Quality

Since the launch of Midjourney in 2022, generative AI has produced over 15 billion images. To put that in perspective, it took professional human photographers and digital artists 150 years to reach that same volume. But this massive output has a dark side: the recursive erosion of quality.

The « Scents of Dali » Dilemma

Imagine prompting an AI to create an image of « a software tester, in a style inspired by Dali. » The AI generates an image, which is eventually uploaded to the web, liked, shared, and ultimately scraped back into future AI training datasets. Over time, the AI is no longer learning from Salvador Dali’s original masterpieces; it is learning from AI-generated approximations of Dali. The result is a watered-down, « faded out » creative substitute.

The AI Feedback Loop: Research shows that AI models age poorly when trained on their own data. This recursive loop leads to « model collapse, » resulting in a measurable decline in performance and output quality—a phenomenon we have already started to observe in major LLMs.

In software testing, this translates to a dangerous dilution of quality. If we rely on AI to generate test cases recursively without strict human oversight, we risk supplanting precise, original test cases with hundreds of degraded, approximate variations, putting approximation at the very heart of quality assurance.

The Silent Rise of « Newspeak » and AI Censorship

To avoid public relations disasters (like Microsoft’s infamous Tay chatbot in 2016), modern LLM creators have implemented strict supervision overlays. While this sounds responsible, the execution raises deep concerns:

  • Exploitative Labor: Much of this content filtering is outsourced to low-wage workers in developing countries, working under the direction of major tech corporations.
  • Anticipated Censorship: If your prompt contains a blacklisted keyword, the AI simply refuses to generate a response, leading to a kind of automated moral policing.
  • Lack of Context: Users attempting to illustrate complex psychological concepts, such as « battered woman syndrome, » often find their requests blocked under the blanket assumption that the topic is too offensive.

This opaque, centralized censorship risks creating a global cultural bias. Because the major LLM developers are concentrated in specific geographic regions, their localized norms and values are projected globally. By restricting the language we can use to prompt these systems, we risk entering a state of AI-driven « Newspeak » where nuance is lost, and complex, sensitive topics cannot even be openly discussed or illustrated.

Conclusion: Why the Human Touch Remains Irreplaceable

As we integrate AI deeper into our professional workflows, one truth remains absolute: human expertise must have the last word.

AI is faster and highly efficient at processing data, but it does not understand our language, our physical nuances, or our emotions. Letting AI operate without strict, skilled human supervision is a recipe for systemic quality degradation. Today, human experts are the only ones capable of separating the wheat from the chaff, ensuring that our reliance on AI does not ultimately destroy the very quality we are trying to protect.


By Olivier Denoo, Vice President of the ISTQB

Laisser un commentaire

Votre adresse e-mail ne sera pas publiée. Les champs obligatoires sont indiqués avec *