Talview Podcast

Agentic AI & Orchestrated Cheating — The Next Threat Wave | Exam Security Summit 2026 | Agentic AI Edition

Talview Season 2 Episode 2

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 38:54

In this podcast episode, you will hear how exam cheating has moved beyond simple tool-assisted misuse into autonomous, agent-driven attacks.

The panel breaks down what orchestrated cheating actually looks like, from unusual timing patterns to behavioral mismatches and multi-candidate signals.

You will also hear which current controls are proving most vulnerable, including lockdown browsers and traditional one-camera proctoring.

The discussion closes with what the next layer of defense should look like, from real-time data forensics to more dynamic assessment design.

This is a sharp, practical conversation on what is already happening now and what exam programs need to prepare for next.

Panel details:

  1. Rory McCorkle (Strasz Assessment Systems)
  2. Charles Mayenga (Fitch Learning)
  3. Steve Kazin (Talview)

Steve Kazan. I'm I'm from Talview. Today's session is about a threat that's already in your programs. Most of you just haven't had a name for it yet. We're gonna spend the next 40 minutes or so giving it a name, understanding what it looks like, figuring out what to do about it. I have two practitioners with me today that I want to welcome. First is Rory McCorkle, VP of Business Development at Stress Assessment Systems. Rory sits at the crossroads of assessment technology and business strategy. He sees this threat from a vendor, excuse me, vendor and a program side at the same time. And our second panelist, Charles Mayenga, Director of Exam Design and Measurement at Fitch Learning. Charles designs and governs high-stakes certification exams for one of the world's leading financial services learning organizations. So welcome to you both. Thank you very much. One quick frame before our first question. There are two very different things people mean when they're talking about AI cheating. The first is a candidate who pastes questions into Chat GPT, a human using a tool. The second is an agentic system, autonomous software that navigates the exam interface, reads the screen, generates the answers, and submits them. No human required. That's not a future scenario. It's really happening now. The problem is that most of our security controls were built to catch the first version. They're not catching the second. And that's pretty much what we're here to talk about today. Let's take a moment and start the moment really by saying that things have really changed because you know I think everyone in this room has a mental model of what AI cheating really looks like. So, Rory, let me start with you. When did you actually first realize that what you were seeing had crossed the line? What was the signal that told you this was really different nowadays? Well, Steve, I'm I'm gonna give away an answer that I'm going to give to a couple of the things that we'll discuss today, and that is data forensics. Um and in fact, uh, you know, I think the first real indication of uh cheating moving from human even proxy test takers who have been a consistent and increasing threat since COVID, um, was when we identified an examination that had been completed uh in literally less than 500 milliseconds. Uh, in fact, some colleagues in the industry uh at Caveon, I was recently at a conference where they identified an examination that had been completed in 180 milliseconds. Now you think about that amount of time, literally, I just took more time saying that sentence than it took for an exam to be completed. Right. Obviously, that is not possible by any sort of human, um, at least not in the current world. So uh that was the clear signal was uh exams being completed in just radically different times with significant accuracy. Awesome. Um Charles, same question to you. When did you first realize that? Yeah, for sure. Just as Rolly has indicated, uh Pastor, thanks for inviting me to join this conversation. It's one of those things that keep us awake, especially when you are handling high-stakes exams and you're looking at what is the implications or what is the consequence of such things happening, uh, because it has both the integrity, both the reputation, as well as the implications to the uh to the organization as well as to the candidates that you serve. Uh, Sorodi has indicated when COVID hit, most of us moved, of course, to online, and that became an accelerator. And all of most organizations, those ones who are even uh not willing or flexible, had no choice, had to move that. So I would say that that's one of the things that moved things to really, really online. And in between there, we have had the development that has happened that has really also accelerated to an extent that we don't even know exactly where it's going to land, but it is really accelerating. One of the things that I did during that particular transition was to, as Rodi has mentioned, to create a forensic approach to looking at what was happening. Because previously, when we had paper copy, we had to look through the hand and writing, we had to look at the notes, we had to go through those kinds of things. Now we were all online. What was the alternative? I came up with an idea of a forensic analysis, a tool that we developed internally and started using it, but realized quickly that that tool is for static forms, fixed forms. So then we realized no, no, no, we need to do more. Especially the challenges were from those online, not those ones who are doing it at the test centers. Most of the challenges have have come from there, and part of what's Rolly has indicated is that you started seeing what we call unusual patterns, unusual behaviors, unusual, because at that point, without evidence, you can't call them cheating. Right. Because it has implications when you can't say that this person was cheating or this person has done that. We generally use the word unusual behaviors, unusual patterns rather than a specific naming of that. And one of them that as Rodi indicated was this unusual completion of nine minutes within a very short time. How is this possible for an exam that is supposed to last three hours? Then we realize that there is more. But the challenge we have now, and we can talk about that later on, is how do you detect that? How do you know that? Because all these things you get to know them after the event. It becomes like you are doing a post-analysis to see that somebody finished nine minutes, somebody are these correct answers very unusually. All these ones are things that you are discovering after the fact. And given that most of the results we provide them online, instant results pass, although it's tentative, but you have already said it pass or successful, and now you are going back to say was it actually a genuine response? Was it a genuine candidate? Was it a genuine attempt that was going on? So one of the things is about that, as Roger S indicated, of completion of uh very short time flame, having answers or responses that match exactly to another person, very you know, and then you find out that these people are actually sitting, were not though they were doing from wherever they were, but the answers seem to match, and then you do this post analysis to find out, and that's when it came uh attention to us that uh there's a lot that is happening here that we don't know, and then the challenge there is that did you stop that or do you proceed or do you find out uh what to do? And we can discuss more about that. That uh that's when it came to us that uh that's a problem. Perfect starting point, Charles. Thank you for that. Um, you know, without naming specific tools, Charles, I'm gonna start with you. Could you walk us through what an orchestrated attack looks like in your uh detection data? You know, what's the signature that distinguishes it from candidates who who are just simply using Chat GPT open in another tab? Yeah, for sure. Without giving the details of some of the things that we do and some of the specific programs, um it's usually to start with the unusual patterns as I've said. When you look at that, and then you realize that this unusual patterns is because of the way in which that particular person, and then if you're looking at some of the videos, the movement between the faces that you see, the hand that you see, and the keyboarding doesn't match, then you realize it says something else more. And then when you examine some of the environment before, although it was proctored and everybody went through that, and then you start seeing some other things that were not visible or were not easy for the proctor to look at. So you look at those unusual patterns, you you look at the unusual things on the wall, things in the room, and then you come to realize that there's more into it than what you are actually uh the proctor will have seen within that short period of time. Right, right, very very true. Hey Rory, I'm gonna uh pass that question on to you now. Um I I don't have too much to add from what Charles has stated. I think um he's hit the nail on the head. Um you either have through data forensics analysis, you either have clear indications with speededness analysis of a non-human actor. Um uh again, humans, we are we are not consistent. Uh, I think yes, uh so uh spending five minutes with any other human can give you that indication. Um technology autonomous systems are consistent, uh typically, and uh unless specifically programmed otherwise. So whether it is extraordinarily rapid response times or for some of the platforms that have been programmed to not create flags for these incredibly, again, impossibly rapid tests, you'll see a very distinct cadence. Um the time per question will match almost perfectly from item to item. So uh speeded in S analysis has been a great indicator. Um, as Charles has also mentioned, you know, certainly looking at uh irregular aberrant consistency between candidates can also give you an indication that uh you have a candidate population that has now moved from individual bad actors to potentially a group of them uh that may be in coordination. Um, I also want to put a big highlight on what Charles has mentioned with regard to movement and behavioral analysis. Um, very true, we've dealt with this for a long time with the proxy test taking problem, wherein someone's movement of their mouse, movement of uh keyboard strokes, what have you, does not match what's occurring on screen. Um, some groups, of course, have developed some great technology to be able to detect that. One of the things that's been uh really critical for us in catching that has been uh a majority of uh groups that we partner with uh using two cameras. Um as you can see here, you know, if I am like this or uh even like this as I normally am when I'm typing, um, you know, it's it's very difficult to identify aberrant movement with that face-on, head-on uh camera analysis. Uh, but uh indeed with that secondary camera, that has been uh great to be able to identify that behavioral physical mismatch in the environment. Great. Thank you for that, Rory. If I could add something about that, um one of the things that uh gives us a clue on the world if you're looking at it, where you have that speededness, you see sometimes situations whereby it takes so long at the beginning, and you wonder and say, how come the person is not making these responses, is not doing all that, and then thereafter you see responses are very a specific pattern. Our realization is that sometimes this is the time that the potential candidate takes time to set up the whole process, and thereafter that things move very quickly because they have set up the system, so that's one indicator they say, Oh, how come this is happening here? And then too at the end, there will be idle time. The idle time is used to convince you or to imply that there's nothing, I'm just waiting to finish time because they don't want to exit the exam situation, the exam environment, so that they can keep up to the end. But at the end is like idle time, but at the beginning, there was also another big idle time, which was actually used for setting it up, and then after that, things work here. The other one I could add is um without still typing any group or any particular uh environment, when you find one, try to find whether there's another one close to that. Most likely, when you find one particular candidate or one particular phenomenon, chances are there is a possibility of another one very in close proximity, whether it is by postal code and institution, workplace, or timing in which they registered their particular exam, uh, or they come back as soon as soon after that. So there is always a coincidence, some kind of it's not just a coincidence that the two of them would be doing that. It is there's a setup that either one of them practices first and doesn't succeed, and the other one executes it, and then the other one follows. So when you see one, don't treat it as an isolated case. Try to find out if there's another one that has got a similar uh pattern. Sometimes they're not very clever. I wish I could advise them. They use the same room, same background, same situation until you feel embarrassed and say, Look, my God, why couldn't you even set it better than this? Exactly. And that's some great insight, Charles. Thank you very much. You know, I I want to ask you both to do the harder thing here. Not what's working, but what you're exposed to. So, my question to you both, uh, let's start with Rory. Um, which of your existing security controls has proven most vulnerable to this type of attack? Um where did you have to say, you know, we built this to catch something different, and it's not really catching this. Is it is it the lockdown browser? Is it behavioral flagging? Rory, could you give us some insights on that? Yeah, Steve. I think first off, it's it's important to note um because uh you know I serve a variety of different groups um through my role, um, knowing your candidate population is key. Um, we have a lot of programs whose candidates, at least heretofore, would never think of this. They're just not in a in a position or in a uh uh point of understanding where they would be using these tools. Um so you know it's important to not throw out the baby with the bathwater, um as it were. Right. But um, all that said, of course, there are candidate populations who are extremely aware um and highly uh uh intelligent about how they're using these tools. Um, I think you you mentioned a couple ones that I would uh uh you know ensure that I would mention that there are more vulnerabilities there. Uh lockdown browser is one of them. Lockdown browser has always been a uh a losing nuclear arms race in the sense that we are always uh, well, uh an organization maintaining a lockdown browser, no matter how rapidly you are measuring what programs candidates have in the background, associating those with uh you know exams that exhibited signs of misconduct and looking for other programs to block, it is a constantly uh reactive battle. Um, and certainly here too, AI has played a role because AI can uh also be very uh smart about how it will relabel or mislabel applications such that they are not blocked, um, down into even editing the registry um such that that works around some of those controls. So, on the one hand, lockdown browser, still very good, very beneficial for a lot of programs, for programs where you do have more sophisticated candidates, um, to a certain extent, its usefulness could be moot. Um, similarly, as I've already mentioned, I would say the traditional one camera head-on remote proctoring methodology has exhibited the most vulnerability, I would say, through a variety of different means, but certainly Agentic has put that um on steroids pickup. That's great. Thanks, Rory. Thanks for that. Um, Charles, how about you? Yes, which of your existing programs? Yeah, yeah. I think I could admire what uh Ruya said, I don't want to dilute that. Uh I would probably say that uh there are those things that at this age that you should actually kind of you cannot do away with them or not even do that. Uh but programs still have those things like fixed reforms, those ones have been um so your questions in your item bank and the secure, those ones are all gone. Uh so it's better to think of better other ways of doing that. Things like item randomization, you know, you randomize the questions, the options, and those kinds of things. Those are sustainable things that you can do. And then you move into what Roll described as multi-layered infrastructure. And I think at Dalview you call it infrastructure. Here you have to have several layers of things that you do that safeguards your program. Are they forensic analysis? Is it how you set up the whole process and what to know? Uh, without deleting what Rod said, one of the things that is important because we are dealing with a very small group of these situations, it's not the large majority of the candidates that are they all have good intentions and they want to do, but there is these few elements who do that. The problem is that when they do that and put your content or whatever they do, the implication is broad, and therefore you uh the consequences are bigger. One thing that I think is important is to have a mechanism for dealing with or expressing the consequences of these behaviors. One thing that I've seen that doesn't succeed is people find all these things and don't find, have no way of making sure that the population that they serve, the people that they work with, the stakeholders know the consequences. When people know that, just like when you're driving on the highway and you know that there's a camera that is going to catch you, you slow down. And when you know that the consequence of that camera is a $500 ticket, you will definitely slow down. Yeah, but most organizations sometimes they catch, they do, but they do nothing, and therefore it continues. But once that's one of the things that I came up with, or we came up with that we have a sequence of events, and one of them is that final consequences. Having an up to the ethics review, up to the final stage, whereby you can say these are the cases, and then use those particular cases that you have caught right at the beginning to whether you're publishing them or you're sharing it, but you are actually saying, look, here are the extreme highway speeding that was at 100 and something that we caught, and this is the amount of fine that they got. So when you've sanctioned those ones, then it helps your program, it protects your program, including the infrastructure that you are building. But if you if you catch and you don't mention, you don't share, you don't have any consequences, uh, this will continue. But once people know that there's a consequence, they're likely to slow down or to be more careful. Right. Steve, I just want to extend because Charles, thank you. You hit such a great point here about just a Decade ago, I was chairing the Association of Test Publishers Security Committee, and we put out a publication. I love one of the things that was very early in the publication was a triangle where we depicted uh in three layers about 60% of the candidates are going to do the right thing. They're going to attempt to use their uh uh uh you know skill sets and follow the rules, all of that good stuff. You know, let's call them 10% of the candidates that tip of the pyramid here are going to cheat no matter what. They're going to find a way. They're fundamentally uh individuals that are committed to that path. What we do have to account for is that group in the middle, and again, we could quibble about percentages, but not the point, um, is is ensuring that those people who would opportunistically take a path of misconduct uh do not. And uh Charles, I so appreciate your comments because uh certainly the ability to call out bad actors, I know not every program has that capability, but at the very least, we've seen when programs clearly delineate, and in more than one place, it is not enough to shove it in your candidate handbook on page 32 out of 48. Um, but emphasizing in various places categories of behavior, example of those, examples of those behaviors, and the repercussions of those behaviors, just that does have a really positive effect at decreasing some of these uh potentials. We're always going to be doing this race on the latest and greatest technology, uh, but you know, outlining repercussions in that way that could indeed have a long-term effect on someone's professional aspirations cannot emphasize that enough. So thank you, Charles. Most welcome. You know, uh those are are real gaps that you both talk about. I suspect a lot of heads are are nodding in this room. Um let's let's let's now flip to the detection side of things. Um, Roy, I wanted to ask you um what behavioral or timing signals have you found to be reliable indicators of agent-driven activity versus a fast, well-prepared human test taker? Is it the response latency? Is it interaction patterns, uh, metadata anomalies? What do you see? Generally, right now, and and you know, I mentioned kind of the categories that have proved effective looking at this. Um, and by the way, I just need to acknowledge my comments here are going to age quickly. Okay. What is true today in this world of particularly agentec AI is likely not going to be true two, three months from now. So, first off, we all have to own that and commit ourselves to continuing to stay up to date. Right now, it has exhibited in one of two forms. It's either these largely incredibly rapid responses uh uh throughout the exams to completion, or it has been consistent patterning, whether that is X number of seconds or milliseconds spent per item, or even a varied pattern that could not be reproduced by a human. So, you know, one second on question one, two on question two, again, some kind of pattern that shows that there's been an attempt to make it past speeded misanalysis, but not sophisticated enough uh to be random. I have no doubt that is going to emerge as the next vector. Um, and uh you know, these programs are going to get more sophisticated at attempting to emulate human behavior uh in a different way. We may also in the future need to look at, for example, when does the uh answer choice get selected versus the next button get hit? That could be an additional step that, by the way, is not measured by most exam delivery drivers at the moment. That could be another area where we could pick up on some of these uh types of behaviors. But fundamentally, it has been that. It has also been looking at what is the parallelism between human movement, human behavior, and what we're seeing on screen, um, which typically is uh, as I mentioned before, with a two-camera view pretty apparent uh on uh on first blush. Again, all of these things with the right sophisticated test taker um and uh evolutions in these platforms can be defeated. So we need to be on the lookout, as Charles said, for the adjacencies. What's the next big leap that these things are going to take? But for now, that's that's where we're at. Great. Thanks, Rory. Charles, same question to you. Do you see any uh of the signals that's happening? Yeah, I I I I do agree with Rory that um those behaviors, those unusual behaviors are indications of what you know, as I said at the beginning, that um when a candidate sort of seem to be taking that amount of time for in the setup, you start the exam and that kind of setting up. And then all of a sudden the speededness happened, those kind of unusual things which are beyond what normally a candidate would would do. Because we used to set up the exam time saying that for this particular exam, the average time a candidate would take is this one and a half minutes or two minutes. Now you see all those unusual uh kind of things, and then you ask yourself, what must be happening here? So that's that's one of the things in terms of time looking at that. Uh, but also, as I said at the beginning, is the environment in which that you are actually uh looking at that particular exam. Is there anything in terms of the collateral environment? Because as you look through, you you see you have to look and see what's the environment looking at. You look at this particular room and say that room is what you expect to be. But then within a short order, then you see a similar room in another particular situation. Some of those things is going to be a coincidence that several rooms have got a similar patterns or similar ways in which they are set up and yet they are far apart or different places. As these things develop, the challenge is how do we catch up? Because they could even be varied depending on which particular places. So we we're doing a catch-up, we're trying to catch up with this kind of change of the technology. But that's where things are going. And therefore, programs have to be ready and start investing more into looking at what else can we look at, not just only the response rate, not only the pattern, but also the environment in which that particular assessment is taking place. Because it's also the environment that that gives you hints or gives you a situation and say that this is what could be uh possibly happening. What is the gray? What is the the how is the background changing, what is happening there, then you will start looking and seeing is that natural or is that uh you know, uh expected to happen, those kind of things. Otherwise, if you focus only on the forensic analysis, we'll still survive, but would not be the only one to look at that you need to, we need to do be very careful about. Right, right. Uh you know, what you what you talk about are the kinds of insights that drive investment decisions. You know, let's close with exactly that. Um, where you would put the next dollar, as they say. So, final question to you both. And if you're taking notes, this is the one to write down. If you had to prioritize one investment or change in your exam security program specifically because of agentic AI, what would it be? Where does the next dollar go? Uh, the next quarter of your team's attention. Where do where does that go? Does it go to detection capabilities, organizational levels, policy and legal? Rory, where do you see that going? Well, I'll I'll differentiate between short and long term, Steve. I think in the short term where we're at right now, um I uh I actually stood up in front of the National College Testing Association conference, I think it was 2017, and I talked about real-time data forensics. At that time, it was a pipe dream. Cloud was still growing, didn't have AI, no longer a pipe dream. Um we haven't cleared that hurdle yet. Um, I I've not seen a uh a program be able to use that at scale quite yet, but that to me is uh a real would be a massive leap in our ability to provide uh proctors and proctoring organizations with uh an incredible toolset through which uh you know these kinds of behaviors which aren't typically apparent to the human invigilator uh and make those obvious. I I think in the long term, um, and again, I don't want to wax too philosophically here, Steve, but I'll I'll leave it as as minute as I can. You know, we need to continue having these conversations about what is the next generation of assessment. Um, the dawn of the internet sparked our last biggest major change, uh uh, you know, which was the move from PBT to CBT. COVID sparked the large-scale expansion of remote proctoring. But at the end of the day, what we have been doing as an industry has not changed largely since the beginning of the 1900s. Uh, you know, the first college board exams came out in early 1900s using multiple choice questions, uh, and we've been doing the same ever since. Uh, I think fundamentally, this is all a uh it's it's potentially challenged, but I think it could be an incredible opportunity for us to look at how we can create uh assessments and assessment processes that can not only be more resistant to these types of tools from a security standpoint, but also improve the fidelity and thus the validity of our assessments. Um, and you know, that could take a number of different forms, and that's not now's not the time, but I'll just say that's something that if your organization is not having an active conversation on how this is affecting you and how it's going to affect your business model in five to ten years, please start having that conversation. Good point. Thanks, Rory. Charles? I like the question. I like the question. Way back, uh I was doing some exams and uh Ron Hamilton, who many people know, said, Well, we give this candidate only four responses, and we try to find out if any of them misbehaved or cheated or something like that. What if we develop 500 of this and tell them that these are the 500 and it's up to them. If they can pick any one of them and respond to that, and then they have done very well. We have covered what it is. Instead of trying to be so secure about these four or five, and we think that we are doing that. Where I'm going with this is that we tend to focus on the output verification. We tend to focus on this small kind of thing. But things are changing. We need to think broaden ourselves and think in a very dynamic way, come up with interactive assessment designs, problem solving, because this is these models are trying to do that. How do we do some of those problem solving situations? Adaptive kind of questioning, because that is what it is that, and you change things. How do we create those kind of environments? Because that's where it is. It's not the traditional way of storing, keeping secure, and then we do we look at the forensic analysis, which is basically the output of what has come out, detected this, is it different from that, what we expected? No. Now we move to another era whereby we don't know what it's going to be. So, how do we start mimicking those kind of things? So, if I was to put money is to encourage and without losing the fidelity, without losing the because you don't want to be testing things which don't exist, which don't practically resemble those. How do you invest into dynamic interactive assessment that will reflect the typical financial situation in out there? The typical financial collusion, collusion. How do you create those kinds of situations? And therefore, because those things are changing. But if you were to come up with examples of testing things that, oh, what if this and this happens? This is what is the answer to this? That's a very uh, you know, and and anybody who answers in a different way probably was cheating. No, things are changing. So we have to invest more in creating dynamic uh assessments and work on that. Together with some of us who work in regulation, how do we make sure that the regulations, the regulatory bodies that we work with are also moving along? Because you can develop all these things, and if the regulations out there are not moving in parallel, then and the regulators are doing something different or they are not caught up with that, then it becomes kind of a problem because it is also what is happening. How do you because at the end of the day, when you say these things, there should be also a regulatory framework that on which will help you make that reference and say when behaviors like this they are outside this regulation, and when these things like this they conform with the existing regulation. So, how do for those of us who are in a regulatory framework, how do you make sure that the regulatory framework also moves along? But TBG is left behind, and in most cases it tends to be a bit traditional, but this time around, it needs to move quickly so that it can actually catch up, but otherwise, it will be all the time chasing. Perfect. You know, uh, both both of you had some great answers to this question. This is exactly the kind of thinking this industry needs more of. Um, with that being said, I want to thank you both, Rory and Charles. Thank you so much. The willingness to be specific, real incidents, real gaps, real priorities that you mentioned is what makes a conversation like this worth having. Three things before I let everybody go to carry out of this room. Um, the threshold has been crossed. You know, agentic cheating is not a future scenario. Programs are encountering it now. The question is, how how do you prepare for it? Uh, the second takeaway, the detection layer has to evolve. You know, static controls at the front door aren't sufficient enough now. The signals that identify agentic behavior live in the continuous session data, behavioral timing, metadata. That's where the next generation of detection lives. And lastly, this is really an organizational challenge, not just a technical one. Programs getting ahead of this, having leadership aligned, and someone who could really translate technical risk into strategic priorities. If that's not you yet, that's something uh first to think of and to fix. So uh again, thank you to the panel. Thank you for everyone joining. Appreciate it. Have a great rest of your day wherever you may be, and um and we'll talk real soon. Thanks again. Thank you.