Md Imanul Huq

Human-centered cybersecurity

Sign in

Research

The least instrumented part of a security system is the person using it.

That is a measurement problem, not an excuse. I treat the person at the keyboard as a subsystem that can be observed with real instruments and characterized under fatigue, deception, and tools that assume too much.

Approach

How I work

Human failure in security is normally inferred after the fact, from an incident report or a click-through rate. That tells you a failure happened; it does not tell you what state the person was in when it happened, or whether the system had already exceeded what any reasonable user could do.

I build end-to-end instrumented studies instead: synchronized EEG and eye-tracking capture alongside behavioral logging, a controlled within-subject design, and statistical treatment appropriate to small samples with many dimensions. The instruments across my studies are a B-Alert X10 wireless EEG system and a GazePoint GP3 HD eye tracker, with analysis pipelines written in Python.

Two commitments run through all of it. First, negative results are reported as negative results. Second, a measured effect is only useful if it points at a design change, so each study ends by asking what the system should do differently.

Line one

Human reliability under adverse conditions

Cognitive fatigue and phishing detection

A within-subject study of 37 participants examining how induced cognitive fatigue changes the ability to distinguish legitimate from fraudulent websites. Correct detection fell sharply after fatigue, and response times got faster rather than slower — a speed–accuracy inversion in which people became quicker and worse at the same time.

The finding I consider most important is a dissociation: participants' own reported fatigue had essentially no relationship to how much their performance had actually degraded. People could not tell they had become unreliable. Fatigue also affected the two directions asymmetrically, impairing the detection of fraudulent sites far more than the recognition of legitimate ones — which is precisely the wrong way round for a defender.

Neurophysiological response to synthetic media

A study of how people classify authentic and deepfake video, combining EEG and gaze with behavioral responses across familiar-face, unfamiliar-face, and look-alike conditions.

The primary group-level analysis found no significant neurophysiological difference between viewing authentic and synthetic video. I report this as a null result rather than reframing it, because it carries a design implication: if the brain does not reliably distinguish the two, detection cannot be delegated to unaided human perception, and the burden has to move into the system.

Line two

When the tool asks too much of the user

Over several semesters I studied how cybersecurity students used PGP encryption in Mozilla Thunderbird. Rather than staging the problem in an artificial usability laboratory, the study grew out of students using the technology as a genuine part of their coursework. As the teaching assistant for that course, I was one of the people they came to when the workflow broke, which is how the recurring failure points became visible in the first place.

The result, published at IEEE PST 2025, identifies the points at which an encryption system defeats users who are technically prepared and motivated. A journal extension is under review.

An ongoing strand extends this to accessibility: OpenPGP communicates trust through visual indicators, and I am examining how those indicators diverge when they are rendered by a screen reader instead of seen. Accessibility here is not a compliance exercise — it is a case where the security signal itself changes depending on how the interface is consumed.

Line three

Evaluating machine reasoning, and human trust in it

I contribute to systematization work that uses large language models as instruments for reasoning over security research corpora — building the retrieval pipelines, the scoring rubrics, and the validation that makes such an evaluation defensible rather than merely automated. In one project this meant a multi-model design with an arbitrator, checked against hundreds of manually coded entries to establish agreement before any conclusion was drawn from the automated pass.

This connects directly to the measurement work. AI-generated explanations can be fluent, confident, and wrong, which makes calibrated reliance — knowing when to trust the output — a human-factors problem of the same shape as phishing detection. Developer over-reliance on AI-generated code is the direction I most want to pursue next.

Publications

Publications

Each entry has a Cite button giving BibTeX, APA, and IEEE formats, ready to copy.

Funding

Funded research participation

My doctoral research was carried out within federally funded projects on which my advisor is principal investigator or subaward principal investigator: AFOSR FA9550-23-1-0453, “Cognitive Security and its Mitigation” (subaward 1564140); NSF CNS-2201465, CNS-2154507, and OAC-2139358; and Department of Defense funded research at Texas A&M, 2024–2025. I participated in these as a graduate research assistant. They are not personal awards to me.

Within them I led neurophysiological user studies on phishing susceptibility and cognitive fatigue, EEG and eye-gaze work on deepfake detection, and studies of email encryption usability.

I am a listed member of the MURI Cognitive Security project team, a multi-university research group directed by Dr. Leanne Hirshfield at the University of Colorado Boulder.