Tuesday, August 11, 2026
HomeSoftware DevelopmentThe AI-augmented tester is right here. So is a brand new drawback:...

The AI-augmented tester is right here. So is a brand new drawback: proving your exams really work

-


Each few years, testing will get rediscovered. I’ve watched this occur greater than as soon as. Years in the past, after I labored at BZ Media, the previous proprietor of SD Occasions, we ran a testing convention and journal, offered them off, sat out a five-year non-compete on the phrase “testing” itself, after which walked again right into a check convention anticipating to discover a modified business. As a substitute we discovered the identical distributors telling the identical tales with the identical instruments. Nothing had moved.

That’s not the world we’re in anymore. In Session 3 of our four-part SD Occasions Reside! Supercast on AI in Testing, “The AI-Augmented Tester: Instruments, Expertise, and Practices,” we talked with Adam Auerbach, Head of Utilized AI for North America at EPAM Methods, about what’s really altering on the bottom. Then LeapWork’s VP of Developer Relations, Donovan Brady, walked by why Playwright has grow to be the default browser automation framework, and what occurs when AI begins writing the exams however no one has time to learn them.

Right here’s what caught with me.

AI doesn’t change what good testing appears like. It raises the stakes for not having it.

Auerbach has been doing this a very long time. Handbook tester, automation architect, QA transformation lead at Capital One, and now the individual EPAM sends in to assist enterprises determine the place AI really matches within the software program supply life cycle. His first level lower towards a number of the AI-testing hype I’ve been listening to all 12 months: the basics haven’t modified. You continue to want traceability, high quality gates, actual check information, and actual environments. What’s modified, he stated, is the price of skipping them.

“I continuously will see organizations who’ve leaned into improvement utilizing AI, after which they’ve important manufacturing points as a result of testing hasn’t stored up,” Auerbach stated. Builders are transport quicker. Testing that was already a bottleneck earlier than AI is now the factor standing between “we moved quick” and “we broke one thing in manufacturing.”

Auerbach laid out a maturity mannequin that’s price internalizing if you happen to’re attempting to determine the place to start out:

  • Degree 1, augmented: AI helps write check circumstances, generate check information, and produce scripts. That is the entry level, and it’s additionally the place most organizations nonetheless are.
  • Degree 2, triage and self-healing: After getting a set of automated exams, AI helps triage failures and, in some circumstances, feed fixes again into the code itself. EPAM’s open-source Report Portal, which now has MCP assist, is constructed for precisely this.
  • Degree 3, spec-driven and agentic: Right here the testing pyramid begins to flip. EPAM’s personal no-code instrument, AgenticQA, sends brokers on to an internet site or cell app to guage what they see and make a judgment name, slightly than executing a set script.

One factor Auerbach was clear about: the talk over whether or not the identical mannequin that writes your code must also check it’s the unsuitable debate. “It’s much less concerning the mannequin and extra about how are you giving the mannequin context about your utility, your small business, the exams that you just’ve run previously, the place you’ve got defects,” he stated. A immediate with no institutional reminiscence behind it would all the time miss issues, no matter which mannequin wrote the code.

And on the tooling shift itself, Auerbach didn’t hedge: Selenium was the pitch for a decade and a half. In the present day, that pitch belongs to Playwright, largely due to how effectively it really works with AI. Which arrange the remainder of the hour completely.

Why Playwright, and why now

Donovan Brady opened with a quantity that’s arduous to argue with: 77 million npm downloads every week, a principally linear climb since Playwright launched in 2020 that turned exponential in roughly the final 12 months. On GitHub, Playwright now sits at almost 100,000 stars, forward of Selenium’s 34,000 and Cypress’s 50,000, regardless of each having a big head begin.

Brady’s framing was helpful: Selenium and Playwright had been by no means actually fixing the identical drawback. Selenium answered “how do I management a browser?” Playwright requested “how do I assist builders ship with confidence?” Selenium was constructed for an internet manufactured from static HTML pages that modified each few months. It assumed browsers behaved persistently and that synchronization points might be dealt with with a well-placed sleep timer. As soon as single-page apps, asynchronous API calls, and steady deployment grew to become the norm, these assumptions broke all of sudden, and testers had been left writing what Brady known as “spaghetti code” simply to guess when a web page was really able to work together with.

Playwright, constructed by former Puppeteer engineers with a clear slate, addressed that instantly: computerized ready, direct browser protocol entry as a substitute of a WebDriver translation layer, remoted browser contexts so one check’s leftover cart objects don’t corrupt the subsequent check’s checkout run, and dramatically higher community mocking and debugging instruments, together with screenshots, video, and time-travel model hint evaluate.

However the actual inflection level, Brady argued, is AI. Playwright remains to be code, and most QA professionals aren’t builders. As soon as AI acquired ok to put in writing strong Playwright scripts on somebody’s behalf, that barrier got here down, and adoption took off.

The enterprise hole: inexperienced doesn’t imply secure

That is the place the session acquired most helpful for anybody operating a testing group slightly than simply writing exams. Brady’s core warning: a passing check tells you a change labored in isolation. It doesn’t let you know it really works throughout your precise enterprise, spanning SAP, Citrix, legacy mainframe techniques, Salesforce, and every little thing else sitting alongside your homegrown functions. “That check that the developer wrote may cross in isolation,” Brady stated. “However what does it appear to be when it’s plugged into all the internet means of all of our functions?”

He additionally shared one thing LeapWork present in its personal home: when the workforce ran Playwright with AI producing the exams, protection numbers appeared nice, 100% passing, full protection. Digging in, they discovered the AI had discovered to put in writing exams that handed, not exams that really validated conduct. No person had time to manually evaluate a whole bunch of AI-generated exams, so no one caught it. Even Playwright itself recommends human verification of AI-written exams, a advice that quietly breaks down the second AI is producing a whole bunch of exams a day.

Brady’s reply to that hole is LeapWork’s new product, LeapWork Play, now in early entry. AI handles authoring, however the ensuing exams run deterministically, with no token price at runtime, evidence-linked reporting, audit logging, and reusable elements groups can share as a substitute of rebuilding exams app by app. It additionally helps MCP, so it could possibly run alongside AI code era instruments like Claude Code, validating adjustments as they’re made slightly than after the actual fact.

If you wish to see it, you may join now at leapwork.ai and get free entry with 1,000 tokens earlier than the product goes totally reside on September thirtieth.

Catch up, and save the date

Should you missed Session 1 (context, brokers, and MCP in distributed testing) or Session 2 (chopping by AI fatigue to deal with outcomes), or if you wish to watch this one, Session 3, in full, together with the complete conversations with Adam Auerbach and Donovan Brady, you may watch the entire Aug. 6 session right here.

The fourth and closing episode of this Supercast collection will air November fifth. If AI is already reshaping how your workforce writes, triages, and validates exams, register now for the Nov. 5 session to avoid wasting your spot.

You may also discover Adam Auerbach’s ongoing writing and commentary on utilized AI in software program supply on LinkedIn.

David RubinsteinDavid Rubinstein

Related articles

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Stay Connected

0FansLike
0FollowersFollow
0FollowersFollow
0SubscribersSubscribe

Latest posts