At 10:40 on a Thursday night, a loan applicant joined a video KYC session at a mid-sized NBFC. Anita, who heads onboarding operations there, wasn’t on the call. She heard about it the next morning from the agent, who was still a little rattled.
The applicant had all the right things. A valid PAN, an Aadhaar-linked address, a clear face that matched the document. He answered the questions calmly. But the agent noticed something she couldn’t quite name. His blinks seemed slightly late, as if the video were a half-beat behind the audio. She asked him to turn his head, then to hold up the PAN card closer to the camera. The image smeared at the edges of his face for a moment, and then the call dropped.
The session was escalated and later flagged as a suspected face-swap attempt. There’s no clean proof of what happened, as these things rarely come with certificates. But Anita told me something I’ve been thinking about since. “For five years we asked, is this the right person? Now we have to ask, is this even a real video?”
That shift, from verifying a person to verifying the session itself, is what I think defines the next phase of remote identity verification. (Anita’s story is an illustrative composite of what ops teams describe, not a single documented case.)
How we got here
Video KYC arrived in India as a practical answer to a simple problem. Branches were far, customers were busy and paper was slow. The Reserve Bank of India’s framework for video-based customer identification, usually called V-CIP, treated a properly conducted video interaction as equivalent to meeting the customer in person. That decision opened the door for banks, NBFCs and fintechs to onboard people across the country without anyone visiting a branch.
The core mechanics are well known by now. A live audio-visual session, a trained official from the regulated entity, a face match against the identity document, liveness checks, geo-tagging, a recorded session and an audit trail. For several years, getting those pieces working reliably was the whole challenge.
Most institutions did get them working. And that’s exactly why the conversation has moved on.
The problem is no longer the camera
When video KYC was new, the main failure modes were mundane. Poor lighting, weak network, an agent queue that kept customers waiting, a document that wouldn’t scan. Drop-offs were high, but the risk was mostly operational.
Today the risk is different. Face-swap tools that work in real time are cheap and widely available. A fraudster doesn’t need to be clever, only to download something. Synthetic identities, built by combining real documents with manufactured details, are harder to catch because every individual piece may look legitimate.
Regulators have noticed. Recent RBI direction has put more weight on liveness and spoof detection, and several industry guides describe deepfake resistance as an expectation now, not a nice extra. The exact technical standards and effective dates differ across sources, so compliance teams should check the current Master Direction and any circulars rather than relying on summaries, including this one.
What matters for the rest of this post is the direction of travel. A basic liveness check that asks someone to blink or smile is no longer a safe assumption. It’s the minimum, and attackers are building around it.
Shift one, from a face to a session
The old model asked whether the face on camera matched the face on the document. The newer model asks a wider set of questions about the entire session.
Is the video feed coming from a real camera or a virtual one? Does the device look the way a genuine customer’s device would look? Is the location consistent with what the customer declared? Did the conversation behave naturally, or were there odd latencies and compression artefacts around the face? Has this device or this face appeared in other applications under other names?
None of these signals is conclusive alone. Together they paint a picture, and that picture is what a good fraud team learns to read. It’s closer to how a seasoned bank manager used to size up a walk-in customer than to a simple pass or fail test.
Shift two, from agents-only to a blend
Under the RBI framework, agent-led V-CIP remains the baseline for banks and NBFCs, and an official has to conduct the interaction. But look at what actually consumes an agent’s day. A lot of it is routine work such as reading out questions, checking that a document is visible, waiting for a customer to find the right card.
The better-run setups are quietly changing this. Software handles the repeatable checks, such as document capture, OCR, face match and liveness scoring, so that the human can focus on judgement. The agent’s job shifts from running a checklist to spotting what doesn’t feel right. Anita’s agent, remember, caught the problem by instinct before any tool did. Good design protects that instinct instead of burying it under a script.
Some regulated contexts allow assisted or self-serve variants with later human review. How far that can go depends on the type of entity and the product, so it’s worth reading the rules for your own category rather than copying a competitor’s flow.
Shift three, from compliance checkbox to customer experience
Here’s something that often gets lost. A video KYC journey is frequently the first time a customer interacts with your brand in a serious way, and it can go badly in ways that have nothing to do with fraud.
Customers drop off when the app asks for permissions they don’t understand, when the call needs a perfect network, when instructions are in a language they don’t speak well, or when they’re told to move to a brighter room at nine at night. People applying for a first loan or a first insurance policy are often less comfortable with technology than your product team assumes.
The next generation of flows seems to take this seriously. Lower-bandwidth video, multilingual prompts, clear guidance before the call starts, and sensible rescheduling when a session fails. These aren’t glamorous, but they decide whether a genuine customer completes onboarding or quietly disappears.
Regulators have also stressed accessibility. Liveness checks shouldn’t depend on specific facial gestures that some customers can’t perform, and there’s an expectation that reasonable accommodations are made. A system that rejects genuine people is failing in a different but equally real way.
Shift four, from one check to a connected journey
Video KYC used to sit by itself. Identity was verified on one screen, documents on another, signing on a third, and risk checks somewhere in a back office.
That’s changing. Institutions are increasingly stitching these steps into a single onboarding journey, where a PAN check, a DigiLocker document fetch, a video session, an e-sign and a risk assessment all feed one decision. The practical benefit is fewer hand-offs and fewer places where a customer, or a fraudster, can slip between steps.
It also changes how you evaluate a provider. The question is less “does your video call work” and more “how does it fit with everything before and after it, and what happens when something fails halfway?”
Privacy is part of the product
Remote identity verification involves recordings of faces, voices, documents and locations. That’s sensitive personal data, and India’s Digital Personal Data Protection Act, 2023 places consent, purpose limitation and security safeguards at the centre of how it should be handled.
In practice, that means telling customers clearly what’s recorded and why, collecting only what the purpose requires, protecting recordings with proper encryption and access controls, and being thoughtful about retention. Sector rules on how long KYC records must be kept differ from general data protection principles, so your legal and compliance teams will need to reconcile the two. I’d be wary of any guide, including blog posts like this one, that gives you a single retention figure with total confidence.
Questions worth asking whatever tools you use
If you’re reviewing your own setup, or choosing a partner, a handful of questions are more useful than any feature list.
How does your system detect a virtual camera or an injected video stream, not just a face on a real camera? Has the liveness technology been independently tested against current deepfake methods? What do agents see when something looks suspicious, and how do they escalate? What happens to a customer whose session fails through no fault of their own? How are recordings stored, who can access them, and how is that access logged? And how quickly can the provider update its models when a new attack pattern appears?
The last question matters more than it looks. Fraud methods change quickly, so a system that was strong two years ago may be weak today unless someone keeps it current.
What I’d expect over the next few years
I’d be cautious about predictions, but a few directions seem fairly safe. Passive liveness and stream-integrity checks will keep growing in importance. Risk signals gathered around the session will matter as much as the video itself. Human reviewers will be used more selectively and with better tools. Journeys will keep getting shorter for low-risk customers and more demanding for higher-risk ones. And regulators will keep adjusting the rules as attack methods evolve, so teams that build for adaptability will have an easier time than those who build for a fixed checklist.
Back to the Thursday night call
Anita’s team changed several things after that session. They added checks for virtual cameras and unusual video behaviour. They gave agents a clearer way to flag a session that feels wrong without needing to prove why. They reviewed how failed calls are followed up. And they started tracking how many genuine customers were being turned away, because a stricter system that rejects honest people isn’t actually safer.
None of it was dramatic. What changed most was how the team thought about the call itself. It stopped being a formality that happened to include a camera and became a decision made with imperfect information, at speed, in a world where seeing is no longer believing.
That, more than any new feature, is what the next generation of remote identity verification asks of the institutions that use it.





Leave a Reply