Claim: Pakistan actor Fahad Mustafa has died, according to video announcements by multiple public figures.
Fact: Fahad Mustafa is alive, contrary to the video statements, which Soch Fact Check found to be AI-manipulated.
In July 2026, multiple social media videos surfaced showing multiple public figures announcing the death of Pakistani actor and TV host Fahad Mustafa. These include journalist Mansoor Ali Khan, Islamic preacher Maulana Tariq Jamil, model Danish Taimoor, and news anchor Waseem Badami.
Many of these videos are accompanied by clickbait captions that Mustafa had died and include clips of the same public figures before a voiceover starts, stating that he was fine.
Additionally, text- and image-based posts also emerged on social media with the same claim; some of these can be viewed here and here.
Fact or Fiction?
Soch Fact Check searched for verifiable news reports about Mustafa’s purported passing by credible media outlets but did not find any.
We also looked at the TV host’s Facebook and Instagram accounts and observed that it is regularly updated.
Reverse-searching the videos of the public figures led us to their original versions, in which there is no mention of Mustafa at all.
For the readers’ convenience, the analysis of each video is organised under a separate subheading, based on the name of the public figure appearing in the clip.
Mansoor Ali Khan
The original video was traced to a vlog by Khan posted on 17 July 2026.
According to Hive Detect’s results, the video is 0.4% likely to be AI-generated with 16.5% chance that it contains AI-generated speech.
DeepFake-O-Meter — a tool developed by the University at Buffalo’s Media Forensics Lab (UB MDFL) — the video is “likely authentic”, with “moderate” confidence. Eight of its available detectors yielded probabilities of 100%, 73.6%, 32.4%, 28.9%, 24%, 23.9%, 6.6%, and 4.7%.
We also ran the video through Global Online Deepfake Detection System (GODDS), a tool developed by Northwestern University’s Security & AI Lab (NSAIL) that uses a combination of various models along with human analysis to provide a holistic summary of the results.
GODDS employed 22 deepfake detection algorithms for the visual content and 70 for the audio component, while two trained analysts also examined the clip.
All predictive models for the visual and audio content said the video “is likely to be fake”:
- The video is likely to be fake with a probability above 0.5, according to seven of the 22 predictive models; it is likely to be fake with a probability below 0.5, according to the 15 other predictive models.
- The audio is likely to be fake with a probability above 0.5, according to 56 of the 70 predictive models; it is likely to be fake with a probability below 0.5, according to the 14 remaining predictive models.
According to GODDS’ human analysts, the video contains “several indicators” that show it may be artificially manipulated. For example:
- The voice does not match the natural cadence and breathing patterns of genuine speech.
- The lip movements do not align with the sounds.
- At the 0:06 mark, a filter appears on Khan’s face.
The analysts added, “The artefacts are largely obscured by poor video quality. Hence, it is difficult to discern possible artefact manipulation from the video’s heavy compression. These indicators can plausibly be explainable by other factors such as motion blur as well.”
The video “is likely manipulated via artificial intelligence”, they concluded.
For further corroboration, we requested an analysis from Shaur Azher, a lecturer who teaches sound design and sound recording at the University of Karachi and the Shaheed Zulfikar Ali Bhutto Institute of Science and Technology (SZABIST). He also works as an audio engineer at our sister organisation, Soch Videos, and specialises in mixing and mastering audio.
Azher explained that for comparison purposes, Sample A is the claim and Sample B is Khan’s actual vlog. He said that based on his calculations and analysis, Sample A is AI-manipulated, or a “synthetic deepfake”, that was created “to simulate a viral celebrity death hoax”.
To support his findings, he provided the following observations:
-
- Mel Frequency Cepstral Coefficients (MFCC) and spectral envelope: In Sample A, the MFCC feature vectors across the vocal pitch range — 100-2.5k hertz (Hz) — exhibit a rigid mathematical uniformity with a variance calculation of only 11.8 decibels (dB), which is symptomatic of a neural vocoder generating synthetic speech. Conversely, Sample B displays an organic cepstral dynamic range of 34.5 dB across complex multi-syllabic geopolitical terms, reflecting authentic physical vocal tract adjustments. The acoustic envelope in Sample A lacks natural micro tremors and smooths over harmonic transitions during emotional inflection points, whereas Sample B demonstrates fluid, biologically consistent spectral energy shifts.
- Phase coherence and stereo field imaging: Cross-channel correlation analysis for Sample A calculates an artificial phase coherence index of 0.99 across both stereo channels, proving it is a digitally synthesised mono file rendered without a physical acoustic space. Sample B exhibits natural inter-channel time differences (ICTD) averaging 0.8 to 1.6 milliseconds, which occurs when sound waves physically interact with a broadcast condenser microphone capsule. This spatial dispersion in Sample B confirms the presence of three-dimensional room reflections in a legitimate broadcast environment, unlike the flat, zero-depth stereo field of Sample A.
- Room-tone fingerprint and noise floor profile: Spectral evaluation of Sample A reveals a sterile background noise floor that abruptly drops to -88 decibels relative to full scale (dBFS) between spoken words, highlighting automated digital gating and a complete absence of ambient room acoustics. Sample B maintains a continuous, authentic newsroom acoustic floor between -55 and -59 dBFS, characterised by low-frequency ambient air and studio ventilation signatures around 50-60 Hz. This continuous acoustic room fingerprint in Sample B proves an active microphone capsule capturing real environmental presence throughout the report.
- Mouth sounds and biomechanical transient analysis: Sample A is entirely devoid of organic biological acoustic markers, showing zero high-frequency transient spikes — 6-9 kilohertz (kHz) where dental friction, tongue movement or labial parting clicks should occur naturally. Furthermore, speech onset in Sample A exhibits a sharp digital fade-in without any preceding respiratory intake or biological breath energy. In Sample B, clear biomechanical micro transients, subtle lip smacks, and pre-speech respiratory inhalations, measuring up to -36 dBFS, are actively detected, verifying authentic physical vocal articulation.
- Full-spectrum frequency band breakdown:
- Sub/low band (20-150 Hz): Sample A exhibits an artificial 8 dB deficit in fundamental chest resonance, whereas Sample B displays robust low end warmth driven by natural broadcast microphone proximity effect.
- Low-mid band (150-500 Hz): Sample A shows algorithmic boxiness and formant averaging around 280 Hz, while Sample B maintains transparent clarity across rapid journalistic delivery.
- High-mid band (2-5 kHz): Sample A exhibits dynamic compression with a narrow 5 dB variance, whereas Sample B projects natural harmonic peaks reaching up to 15 dB during vocal emphasis.
- High-frequency band (8-20 kHz): Sample A displays a steep digital low-pass filter roll off at approximately 11 kHz typical of voice-cloning synthesis models, whereas Sample B extends cleanly past 16 kHz with organic sibilant decay.
Maulana Tariq Jamil
Soch Fact Check found a matching visual of Jamil in a YouTube video uploaded on 24 January 2023, indicating that the clip the screenshot was taken from predates the viral one.
According to Hive Detect’s results, the video is 0% likely to be AI-generated with a 1.3% chance that it contains AI-generated speech.
DeepFake-O-Meter said the video turned up “mixed signals”, with “low” confidence. Eight of its available detectors yielded probabilities of 0.3%, 2%, 30.1%, 97.2%, 100%, 0.5%, 33.2%, and 62.5%.
On the other hand, GODDS employed 22 deepfake detection algorithms for the visual content and 70 for the audio component, while two trained analysts also examined the clip.
All predictive models for the visual and audio content said the video “is likely to be fake”:
- The video is likely to be fake with a probability above 0.5, according to 15 of the 22 predictive models; it is likely to be fake with a probability below 0.5, according to the seven other predictive models.
- The audio is likely to be fake with a probability above 0.5, according to 38 of the 70 predictive models; it is likely to be fake with a probability below 0.5, according to the 32 remaining predictive models.
According to GODDS’ human analysts, the video contains “several indicators” that show it may be artificially manipulated. For example:
- The voice does not match the natural cadence and breathing patterns of genuine speech.
- The lip movements do not align with the sounds.
- Small dot-like artefacts are visible throughout the video. While they may be consistent with a display overlay, they could also indicate manipulation.
- A partially cropped logo is visible at the top of the frame and rotates throughout the video, indicating that the footage may have been captured from another platform.
The analysts further said, “The artefacts are largely obscured by poor video quality. Hence, it is difficult to discern possible artefact manipulation from the video’s heavy compression. These indicators can plausibly be explainable by other factors such as motion blur as well.”
The video “is likely manipulated via artificial intelligence”, they concluded.
Since the original video could not be identified, we provided Azher with another one of Jamil’s videos — from 11 July 2026 — with a similar environment.
He explained that for comparison purposes, Sample A is the claim and Sample B is the video of Jamil that we provided. He said that based on his calculations and analysis, Sample A is AI-manipulated — or a “synthetic deepfake” — and Sample B is “human”.
To support his findings, he provided the following observations:
- MFCC and spectral envelope: Sample A exhibits hyper-smooth, mathematically-rigid MFCC trajectories that lack the micro transient variations typical of a biological vocal tract. Neural vocoders used in voice cloning often average out these complex spectral envelopes, resulting in an unnaturally consistent harmonic distribution. In contrast, Sample B displays organic, non-linear shifts across lower and upper cepstral coefficients that accurately reflect real-time physical adjustments in the speaker’s vocal cords and mouth shape.
- Phase coherence and stereo field imaging: Analysis of the stereo field in Sample A reveals an artificially-locked, phase-coherent centre image with zero spatial drift or authentic ICTD. Synthetic speech generation typically renders a dry, mathematically-centred mono signal that lacks the complex phase relationships created by physical sound waves interacting with a room. Sample B demonstrates natural phase dispersion and subtle inter-channel micro-delays, proving the acoustic presence of a physical microphone capsule capturing sound in a three-dimensional space.
- Room tone fingerprint and noise floor profile: Sample A presents a sterile, completely deadened noise floor devoid of any authentic acoustic room tone or environmental acoustic reflections. as clean audio is actually a synthetic vacuum where background energy drops to absolute zero between syllables, a common artefact of AI gating and synthesis. Sample B contains a continuous, distinct acoustic room fingerprint with natural low-frequency room modes and ambient air movement that seamlessly integrates with the speaker’s voice.
- Mouth sounds and biomechanical transient analysis: Sample A is entirely devoid of biological micro transients, showing no evidence of saliva clicks, lip smacks, tongue movements or natural respiratory intakes. Neural text-to-speech (TTS) models consistently fail to generate these random, high-frequency biological imperfections without sounding distorted. Sample B prominently features authentic biomechanical clicks, crackles, and natural breathing pauses that occur organically between words and during syllabic transitions.
- Full spectrum frequency band breakdown: The low-frequency bands (20-250 Hz) in Sample A are noticeably deficient in natural sub harmonic warmth, while the mid range (250-4,000 Hz) is heavily restrained by what behaves like an aggressive synthetic multiband compression. In contrast, Sample B exhibits strong, dynamic low-end energy driven by natural microphone proximity effect and an uncompressed, direct mid range. The high frequencies (4 kHz+) in Sample A taper off with a synthetic, sterile smoothness, whereas Sample B retains the natural air and frictional sibilance of human speech.
Danish Taimoor
The original video featuring Taimoor was posted on 21 March 2026.
According to Hive Detect’s results, the video is 0.1% likely to be AI-generated with a 5.5% chance that it contains AI-generated speech.
However, between the 0:11 and 0:14 marks, the probability of the clip containing AI-generated speech rises to 67.1%.
DeepFake-O-Meter returned an “inconclusive” result, with “very low” confidence. Eight of its available detectors yielded probabilities of 99.9%, 98.2%, 45%, 93.9%, 6.3%, 99.8%, 26.3%, and 1.2%.
GODDS employed 22 deepfake detection algorithms for the visual content and 70 for the audio component, while two trained analysts also examined the clip.
All predictive models for the visual and audio content said the video “is likely to be fake”:
- The video is likely to be fake with a probability above 0.5, according to nine of the 22 predictive models; it is likely to be fake with a probability below 0.5, according to the 13 other predictive models.
- The audio is likely to be fake with a probability above 0.5, according to 63 of the 70 predictive models; it is likely to be fake with a probability below 0.5, according to the seven remaining predictive models.
According to GODDS’ human analysts, the video contains “several indicators” that show it may be artificially manipulated. For example:
- The voice does not match the natural cadence and breathing patterns of genuine speech.
- The lip movements do not align with the sounds.
- The video contains the logo of Green TV Entertainment, which does not align with the claimed source of the footage. This indicates that the content may have been repurposed from another source.
“The artefacts are largely obscured by poor video quality. Hence, it is difficult to discern possible artefact manipulation from the video’s heavy compression. These indicators can plausibly be explainable by other factors such as motion blur as well,” the analysts said.
According to them, the source is a 20 March 2026 YouTube video, which is actually the promo for the programme we identified that was uploaded on 21 March 2026.
The video “is likely manipulated via artificial intelligence”, the analysts concluded.
Azher offered the same conclusion that Sample A is a “synthetic deepfake” — meaning it was manipulated using AI tools — that was “generated to simulate a viral celebrity death announcement”, whereas Sample B is human.
To back up his conclusion, he provided the following observations:
-
- MFCC and spectral envelope: In Sample A, the lower-order MFCCs exhibit a mathematically-constrained variance of only 10.4 dB across the vocal fundamental frequency range — 120-2,200 Hz — strongly indicating algorithmic smoothing from a neural TTS vocoder. Conversely, Sample B demonstrates an organic, dynamic cepstral range of 35.8 dB across rapid polysyllabic news delivery, reflecting authentic physical vocal tract and glottal adjustments. The spectral envelope in Sample A lacks natural harmonic micro tremors and exhibits rigid transition slopes during emotional inflection points, whereas Sample B displays fluid, biologically-consistent spectral energy shifts.
- Phase coherence and stereo field imaging: Cross-channel phase correlation calculations for Sample A yield an artificially-locked coherence coefficient of 0.98 across both stereo channels, confirming a digitally-synthesised mono audio file rendered without physical acoustic space dispersion. In contrast, Sample B exhibits natural ICTD averaging between 0.7 to 1.5 milliseconds, which occurs when sound waves physically interact with a broadcast condenser microphone capsule in a real room. This spatial dispersion and phase complexity in Sample B confirm authentic three-dimensional room reflections and microphone diaphragm behavior that cannot be replicated by standard voice-cloning vocoders.
- Room tone fingerprint and noise floor profile: Spectral measurement of the background noise floor in Sample A shows an abrupt, sterile drop to -89 dBFS during pauses between words, highlighting algorithmic gating and a complete absence of continuous room acoustics or air presence. Sample B maintains a consistent broadcast studio ambient floor between -53 and -57 dBFS, featuring natural low-frequency heating, ventilation, and air conditioning (HVAC) and room acoustic signatures around 45-55 Hz. This continuous, unbroken acoustic room fingerprint in Sample B proves an active broadcast microphone capturing real environmental presence throughout the video, unlike the synthetic vacuum of Sample A.
- Mouth sounds and biomechanical transient analysis: Sample A is completely devoid of organic biological acoustic markers, showing zero high-frequency transient energy spikes — 6.5-9.5 kHz — where dental friction, labial parting clicks or tongue movements should naturally occur during rapid speech. Furthermore, syllabic onset in Sample A displays a sharp digital fade-in without any preceding biological breath or respiratory intake energy. In Sample B, clear biomechanical micro transients, subtle lip smacks, and pre-speech respiratory inhalations measuring up to -35 dBFS are actively recorded, verifying organic human vocal articulation and physical vocal cord engagement.
- Full-spectrum frequency band breakdown:
- Sub/low band (20-150 Hz): Sample A exhibits an artificial 9 dB deficit in fundamental chest resonance and acoustic proximity effect, whereas Sample B displays robust low-end warmth characteristic of a broadcast microphone capture.
- Low-mid band (150-500 Hz): Sample A shows algorithmic boxiness and formant resonance peaking around 310 Hz, while Sample B maintains transparent, unhyped clarity across rapid word transitions.
- High-mid band (2-5 kHz): Sample A demonstrates dynamic flattening with a compressed 4.5 dB variance, whereas Sample B projects natural harmonic inflection peaks, reaching up to 16 dB during vocal emphasis.
- High-frequency band (8-20 kHz): Sample A displays a steep digital low-pass filter roll-off at approximately 10.8 kHz typical of synthetic speech generation models, whereas Sample B extends cleanly above 16 kHz with organic sibilant decay.
Waseem Badami
Soch Fact Check traced the original video to YouTube, where it was uploaded on 19 February 2026.
According to Hive Detect’s results, the video is 0% likely to be AI-generated with 0% chance that it contains AI-generated speech.
DeepFake-O-Meter said there were “mixed signals”, with “low” confidence. Seven of its available detectors yielded probabilities of 100%, 96.9%, 60.2%, 41%, 99.5%, 98.9%, and 2.3%.
GODDS employed 22 deepfake detection algorithms for the visual content and 70 for the audio component, while two trained analysts also examined the clip.
All predictive models for the visual and audio content said the video “is likely to be fake”:
- The video is likely to be fake with a probability above 0.5, according to 18 of the 22 predictive models; it is likely to be fake with a probability below 0.5, according to the four other predictive models.
- The audio is likely to be fake with a probability above 0.5, according to 28 of the 70 predictive models; it is likely to be fake with a probability below 0.5, according to the 42 remaining predictive models.
According to GODDS’ human analysts, the video contains “several indicators” that show it may be artificially manipulated. For example:
- The voice does not match the natural cadence and breathing patterns of genuine speech.
- The lip movements do not align with the sounds.
According to them, the source is a 19 February 2026 YouTube video, which is from the same set and date as the one we identified. The former was posted by ARY Digital HD, whereas the latter by Shan e Ramazan, the dedicated channel for the media outlet’s special Ramzan transmission called “Shan-e-Ramzan.”
The analysts observed, ““The artefacts are largely obscured by poor video quality. Hence, it is difficult to discern possible artefact manipulation from the video’s heavy compression. These indicators can plausibly be explainable by other factors such as motion blur as well.”
The video “is likely manipulated via artificial intelligence”, they concluded.
Noting, again, that Sample A is manipulated using AI tools — or a “synthetic deepfake” — and Sample B is “human”, Azher provided the following observations to back up his conclusion:
-
- MFCC and spectral envelope: In Sample A, the lower-order MFCC variance across Badami’s signature nasal formant range (1.2-2.8 kHz) calculates to an unnaturally-compressed 14 dB variance, indicating a neural vocoder smoothing out rapid speech transitions. Conversely, Sample B exhibits an organic cepstral dynamic range of 32 dB, reflecting real-time physical vocal tract adjustments during his rapid speech delivery. The acoustic envelope in Sample A demonstrates algorithmic formant averaging that fails to capture the abrupt pitch inflection milestones present in authentic human articulation.
- Phase coherence and stereo field imaging: Inter-channel correlation calculations for Sample A reveal a mathematically-locked phase coherence coefficient of 1.0 across left and right channels, proving a synthetic mono rendering without acoustic space dispersion. Sample B displays natural ICTD fluctuating between 0.5 and 1.4 milliseconds, caused by subtle physical head movement relative to a broadcast condenser microphone. The spatial imaging in Sample B confirms authentic three-dimensional room reflections that cannot be replicated by standard TTS vocoders.
- Room tone fingerprint and noise floor profile: Spectral measurement of the background noise floor in Sample A shows an abrupt digital drop to -86 dBFS between spoken words, indicating synthetic silence and algorithmic gating. Sample B maintains a consistent broadcast studio ambient noise floor between -54 and -58 dBFS, featuring natural low-frequency HVAC room modes around 60 Hz. This continuous acoustic room fingerprint in Sample B confirms an open-studio microphone capturing physical environmental air throughout the recording.
- Mouth sounds and biomechanical transient analysis: Sample A completely lacks biological acoustic markers, showing zero high-frequency transient spikes — between 5-8 kHz — where labial parting clicks or dental friction should naturally occur during Badami’s fast delivery. Furthermore, speech onset in Sample A displays a digital fade in without any pre-speech respiratory intake energy. In Sample B, clear biomechanical micro transients, audible sighs, and pre-speech respiratory inhalations measuring up to -38 dBFS are present, confirming organic physical speech production.
- Full-spectrum frequency band breakdown:
- Sub/low band (20-150 Hz): Sample A lacks the fundamental chest resonance and acoustic proximity effect — a -6 dB deficit — present in Sample B’s broadcast microphone capture.
- Low-mid band (150-500 Hz): Sample A shows artificial algorithmic resonance peaking around 320 Hz, whereas Sample B remains transparent and naturally articulated across rapid-word transitions.
- High-mid band (2-5 kHz): Sample A exhibits dynamic flattening with a compressed 6 dB range, while Sample B displays natural projection peaks up to 14 dB during vocal emphasis.
- High-frequency band (8-20 kHz): Sample A demonstrates a steep digital low-pass filter roll-off at approximately 11.5 kHz typical of 22.05 kHz voice-training models, whereas Sample B extends cleanly above 16 kHz with natural sibilant decay.
Separately, we traced the videos to a TikTok account that appears to be the source of the fake news. It has since changed its handle from @2jdjfjcsv6s to @syatt480.
@syatt480 posted one video featuring Mansoor Ali Khan, which was viewed over 3.5 million times. It shared two clips showing Maulana Tariq Jamil that gained more than 1.9 and 5.3 million views.
The same account posted one video featuring Danish Taimoor that was viewed over 1.7 million times and three separate clips of Waseem Badami that garnered more than 593,100, 556,100, and 483,600 views so far.
The TikTok user @syatt480 has previously spread false claims about Pakistani cricketer Wasim Akram’s death that we have already debunked.
Soch Fact Check, therefore, concludes that the viral videos are AI-manipulated and that Fahad Mustafa is alive.
Virality
Soch Fact Check found the AI-manipulated videos shared multiple times on Facebook and Instagram.
The claim, sans the manipulated clips and public figures, also spread to YouTube and was shared in standalone Facebook posts as well.
Conclusion: Fahad Mustafa is alive, contrary to the video statements, which Soch Fact Check found to be AI-manipulated.
Cover photo: Soch Fact Check / Fahad Mustafa
To appeal against our fact-check, please send an email to appeals@sochfactcheck.com