Claim: Different videos show Bollywood actor Akshay Kumar speaking in favour of Imran Khan and asking people to raise their voice for Pakistan’s former PM, while expressing love for him and condemning the way he was arrested.
Fact: All four videos have been manipulated using artificial intelligence (AI) tools; none of the original clips show Kumar speaking in favour of Khan.
In August 2026, four videos of Bollywood actor Akshay Kumar speaking in favour of the Pakistan Tehreek-e-Insaf’s (PTI) incarcerated founder and former Pakistani Prime Minister, Imran Khan, went viral on social media.
The transcription of his remarks from the four videos are as follows:
- Video 1: “You also need to raise your voice. This always happens to good people. Everyone has to unite for Imran Khan and I’m supporting Imran Khan.”
- Video 2: “What is this? Is this the way to arrest the president? In this manner? Like animals? I’m supporting Imran Khan, who was the president but my love for him is still there in my heart.”
- Video 3: “Hi guys, this is Akshay Kumar. First of all, watch this video. Walaikum Salam. How are you? Imran Khan the best leader of the world, welcome to TikTok. Have you followed him or not? Follow him. Welcome Imran Khan. Love you.”
- Video 4: “Tell me in these 24 hours how much love do you have for Imran Khan? I have a lot of love [for him]. Akshay Kumar love Imran Khan. Now fill up the comment [section].”
The ex-PM remains imprisoned over various legal cases since August 2023.
Fact or Fiction?
Soch Fact Check searched for credible news reports by reputable media outlets on the Bollywood actor speaking in favour of Khan, but did not find any.
Reverse-searching all four of the viral clips led us to their original versions, in which there is no mention of the former PM at all, indicating that they may have been manipulated using AI tools.
The first and second viral clips are manipulated versions of an authentic video from 22 November 2019, when the Bollywood actor promoted a device manufactured by the Indian fitness technology company GOQii.
The third viral video is a doctored version of Kumar announcing on 3 September 2022 the 14th Akshay Kumar International Kudo Tournament, which took place from 24 to 26 October 2022 in Bardoli, a town in Gujarat, India.
The fourth viral clip is an altered version of the actor’s 5 January 2017 video condemning what came to be known as a “mass molestation” event in Bangalore, India, on New Year’s Eve. The assault was reported by multiple news outlets.
All videos were then run through various deepfake detectors and an audio engineer has also provided his analysis on them. For the readers’ convenience, assessments of the aforementioned individual clips have been organised under separate subheadings.
Video 1
Audio engineer’s analysis
To corroborate our suspicions about AI manipulation, we sought a comment from Shaur Azher, a lecturer who teaches sound design and sound recording at the University of Karachi and the Shaheed Zulfikar Ali Bhutto Institute of Science and Technology (SZABIST). He also works as an audio engineer at our sister organisation, Soch Videos, and specialises in mixing and mastering audio.
Azher explained that for comparison purposes, Sample A is the claim and Sample B is Kumar’s video from 22 November 2019. “When examining the acoustic behaviours, the evidence points to entirely different recording and processing conditions for the two files,” he said.
“The brighter spectrum, layered artificial ambience, frequent transient artefacts, and compressed dynamic range all indicate that Sample A is heavily-processed and consistent with added artificial layers. Conversely, Sample B maintains the naturally embedded room tone and typical characteristics of an authentic real world environment. However, it should be noted that the spoken topics in the two clips are different, meaning they do not stem from the same underlying recording,” he added.
To support his findings, he provided the following observations:
- Voice fingerprint: When the underlying mathematical structure of both audio clips is analysed, the overall Mel-Frequency Cepstral Coefficients (MFCC) profile shows distinct differences in the spectral envelope and timbre between the two recordings. Furthermore, Sample A displays a higher mean and variance in spectral flux, indicating more rapid spectral shifts over time, compared to Sample B.
- Dynamic range and compression: Sample A has a lower dynamic range, with a crest factor of approximately 6.3, making it appear much more compressed. In contrast, Sample B has a crest factor of roughly 10.1, showing a higher dynamic range that is typical of a natural, unprocessed recording.
- Mouth clicks and processing artefacts: Sample A actually exhibits a higher overall activity in its zero crossing rate and a greater high-frequency energy ratio than the original. Sample A contains more frequent high-frequency transient bursts, around 1.36 per second, compared to Sample B’s 0.86 per second, which is consistent with processing artefacts or unnatural mouth noises.
- Fake background static: The noise floor in Sample B is heavily low-frequency dominant, consisting of 96% low energy, which is consistent with the natural ambient rumble of a real physical room. Sample A, however, features a noise floor with substantial mid-frequency energy — about 57% — and reveals a more uniform, layered horizontal structure in quieter regions. This indicates that Sample A has an added or artificial ambient noise bed rather than naturally embedded room tone.
- Unnatural pitch: Sample A exhibits a brighter spectrum with more elevated mid- and high-frequency energy, possessing a spectral centroid of approximately 2,280 hertz (Hz). By comparison, Sample B has a darker spectral balance with a spectral centroid of roughly 1,425 Hz, retaining a much stronger and more natural low-end bass presence.
Deepfake detectors’ results
We tested two portions of the audio component in Hiya Deepfake Voice Detector: “This always happens to good people” and “Everyone has to unite for Imran Khan and I’m…”. The tool said the sampled voice is “likely authentic” for both, with scores of 99 out of 100 each.
According to results from Hive Detect, the media is 0% likely to be AI-generated video, 0% likely to have AI-generated speech, and 0% likely to be deepfake.
DeepFake-O-Meter, a tool developed by the University at Buffalo’s Media Forensics Lab (UB MDFL), found the video to contain “mixed signals”, with “low” confidence. Six of its detectors gave probabilities of 99.9%, 41.4%, 71.1%, 46.8%, 97.3%, and 98.1%.
On the other hand, NVIDIA Synthetic Video Detector said the clip is “synthetic”, with a score of 92.3%, which is “greater than threshold (30%), meaning it’s likely synthetic or AI-generated”.
Moreover, Global Online Deepfake Detection System (GODDS) — a tool developed by Northwestern University’s Security & AI Lab (NSAIL) that uses a combination of various models along with human analysis to provide a holistic summary of the results — also found the video to be tampered with using AI.
The tool employed 22 deepfake detection algorithms for the visual content and 70 for the audio component, while two trained analysts also examined the clip.
All predictive models for the visual and audio content said the video “is likely to be fake”:
- The video is likely to be fake with a probability above 0.5, according to 14 of the 22 predictive models; it is likely to be fake with a probability below 0.5, according to the eight other predictive models.
- The audio is likely to be fake with a probability above 0.5, according to 52 of the 70 predictive models; it is likely to be fake with a probability below 0.5, according to the 18 remaining predictive models.
According to GODDS’ human analysts, the video contains “several indicators” that show it may be artificially manipulated. For example:
- Between the 0:00 and 0:02 marks, the left eyebrow appears to blend in with the eye, and the eyelid is not visible.
- There is an unidentifiable man in the top region of the video.
- The lip movements do not align with the sounds.
“The artefacts are largely obscured by poor video quality. Hence, it is difficult to discern possible artefact manipulation from the video’s heavy compression. These indicators can plausibly be explainable by other factors such as motion blur as well,” the analysts said, concluding that “this media is likely manipulated” using AI.
Video 2
Audio engineer’s analysis
For comparison purposes, Sample A is the claim and Sample B is Kumar’s video from 22 November 2019. Azher explained that when the acoustic behaviours of the two clips are examined, “the evidence points to different recording conditions and processing for the two files”.
The audio engineer said, “The brighter spectrum, layered artificial ambience, shorter decay time, frequent transient artefacts, and compressed dynamic range all indicate that Sample A is heavily-processed and consistent with added ambient layers.
“Conversely, Sample B maintains a natural low-frequency dominant room tone and longer natural decay consistent with a real space. It is noteworthy that the spoken content in the two clips is entirely different, confirming they are not the same underlying source recording,” he added.
To support his conclusion, Azher provided the following observations:
- Voice fingerprint: When the underlying mathematical structure of both audio clips is analysed, the overall MFCC profile shows a clear divergence in the spectral envelope and timbre between the two recordings. Furthermore, Sample A displays a higher mean and variance in spectral flux, indicating a more rapid spectral movement over time, compared to Sample B.
- Dynamic range and compression: Sample A has a lower dynamic range with a crest factor of approximately 4.9, making it noticeably more compressed and processed. In contrast, Sample B has a crest factor of roughly 10.1, showing a higher dynamic range that is typical of a natural recording.
- Mouth clicks and processing artefacts: Sample A exhibits a higher overall activity in its zero crossing rate than the original. Although their high-frequency energy ratios are nearly identical, Sample A contains more frequent short bursts and high-frequency transient events, around 1.40 per second, compared to Sample B’s 1.03 per second, which is consistent with processing artefacts or unnatural mouth noises.
- Fake background static: The noise floor in Sample B is heavily low-frequency dominant, consisting of 95% low energy, which is consistent with classic natural room tone. Sample A, however, features a higher overall noise floor with significant mid-frequency content — about 28% — and reveals more visible structure in quieter regions. This indicates that Sample A is more consistent with an added or artificial ambient noise bed rather than naturally embedded room tone.
- Unnatural pitch: Sample A exhibits a brighter spectrum overall, possessing a spectral centroid of approximately 1,870 Hz. By comparison, Sample B has a spectral centroid of roughly 1,425 Hz, retaining a much stronger and natural low-mid weight.
Deepfake detectors’ results
Soch Fact Check tested two portions of the audio in Hiya Deepfake Voice Detector: “the president? In this manner? Like animals?” and “I’m supporting Imran Khan, who was the president but my love…”. According to the tool, its “models didn’t detect enough voice” for the first one, while the second one “is likely authentic”, it said, with a score of 95 out of 100.
Hive Detect’s results said the media is 0.3% likely to be AI-generated video, 0% likely to have AI-generated speech, and 0% likely to be deepfake.
According to DeepFake-O-Meter, there were “mixed signals”, with “low” confidence. The six detectors we used provided probabilities of 100%, 22.2%, 33.8%, 79.6%, 91.9%, and 88.3%.
On the other hand, NVIDIA Synthetic Video Detector said the clip is “synthetic”, with a score of 75.2%, which is “greater than threshold (30%), meaning it’s likely synthetic or AI-generated”.
Moreover, GODDS — which employed 22 deepfake detection algorithms for the visual content and 70 for the audio component, while two trained analysts also examined the clip — said all predictive models the video “is likely to be fake”:
- The video is likely to be fake with a probability above 0.5, according to 14 of the 22 predictive models; it is likely to be fake with a probability below 0.5, according to the eight other predictive models.
- The audio is likely to be fake with a probability above 0.5, according to 52 of the 70 predictive models; it is likely to be fake with a probability below 0.5, according to the 18 remaining predictive models.
According to GODDS’ human analysts, the video contains “several indicators” that show it may be artificially manipulated. For example:
- At the 0:00 and 0:12 marks, the eyebrows appear to blend in with the eyes and the eyelids are not visible.
- The lip movements do not align with the sounds.
“The artefacts are largely obscured by poor video quality. Hence, it is difficult to discern possible artefact manipulation from the video’s heavy compression. These indicators can plausibly be explainable by other factors such as motion blur as well,” the analysts said, adding that they believe the clip “is likely manipulated” via AI.
Video 3
Audio engineer’s analysis
For comparison purposes, Sample A is the claim and Sample B is Kumar’s video from 3 September 2022. Azher explained that examining the acoustic behaviours “points to different recording conditions and processing for the two files”.
“The brighter spectrum, layered artificial ambience, shorter decay time, frequent transient artefacts, and compressed dynamic range all indicate that Sample A is heavily processed and consistent with added ambient layers, which could prove that the text-to-speech (TTS) AI generation method was used.
“Conversely, Sample B maintains a natural low-frequency dominant room tone and longer natural decay consistent with a real space. It may be noted that the spoken content in the two clips is entirely different, confirming they are not the same underlying source recording,” he said.
To support his findings, Azher provided the following observations:
- Voice fingerprint: Sample A displays a higher mean (0.046) and standard deviation (0.034) in spectral flux, compared with those of Sample B (0.025 and 0.025, respectively), indicating a more rapid spectral movement over time.
- Dynamic range and compression: Sample A has a lower dynamic range with a crest factor of approximately 5.5, making it noticeably more compressed and processed. In contrast, Sample B has a crest factor of roughly 10.9, showing a higher dynamic range that is typical of a natural recording.
- Mouth clicks and processing artefacts: The zero-crossing-rate (ZCR) means are nearly identical (0.045 and 0.046 for Samples A and B, respectively), as are the high-frequency (>5 kHz) energy ratios — greater than 5 kilohertz (kHz) — at 0.029 and 0.036). The detected click-like / high-frequency transient events occur at similar rates (1.62 and 1.79 per second, respectively), so there is no strong evidence that one clip contains markedly more mouth click artefacts than the other.
- Fake background static: The noise floor in Sample B is heavily low-frequency dominant, which is consistent with classic natural room tone. Sample A, however, features a higher overall noise floor with substantial mid-frequency content and more structured quieter regions. This indicates that Sample A is more consistent with an added or artificial ambient noise bed rather than naturally embedded room tone.
- Spectral balance / brightness: Sample A is heavily-weighted in the low-mid band, while Sample B shows stronger true mid-range energy and higher high-frequency content, resulting in a brighter overall spectrum.
- Artificial reverb / decay: Envelope autocorrelation decay — which is the time to 0.1 seconds — is shorter for Sample A at 219 milliseconds (ms) than for Sample B, which is 350 ms. Sample B exhibits a longer, more natural-sounding decay; Sample A decays faster (drier or differently processed space).
Deepfake detectors’ results
We used two phrases in Hiya Deepfake Voice Detector: “Hi guys, this is Akshay Kumar. First of all, watch this video.” and “Imran Khan, the best leader of the world, welcome to TikTok.” It said both samples were “likely authentic”, with scores of 99 and 78 out of 100.
Hive Detect’s results revealed that the sample is 0% likely to be AI-generated video, 0% likely to have AI-generated speech, and 0% likely to be deepfake.
DeepFake-O-Meter said its calculations were “inconclusive”, with “very low” confidence. The six detectors we used provided probabilities of 1.1%, 23.3%, 98%, 46.7%, 51.7%, and 2.5%.
On the other hand, NVIDIA Synthetic Video Detector found the video to be “synthetic”, with a score of 80.1%, which is “greater than threshold (30%), meaning it’s likely synthetic or AI-generated”.
Moreover, GODDS also found the video to be AI-manipulated. It employed 22 deepfake detection algorithms for the visual content and 63 for the audio component, while two trained analysts also examined the clip.
All predictive models for the visual and audio content said the video “is likely to be fake”:
- The video is likely to be fake with a probability above 0.5, according to five of the 22 predictive models; it is likely to be fake with a probability below 0.5, according to the 17 other predictive models.
- The audio is likely to be fake with a probability above 0.5, according to 57 of the 70 predictive models; it is likely to be fake with a probability below 0.5, according to the six remaining predictive models.
According to GODDS’ human analysts, the video contains “several indicators” that show it may be artificially manipulated. For example:
- At the 0:01 mark, the finger appears to blend into the nose.
- At the 0:05 mark, the hand appears to blur outwards.
- The lip movements do not align with the sounds.
“The artefacts are largely obscured by poor video quality. Hence, it is difficult to discern possible artefact manipulation from the video’s heavy compression. These indicators can plausibly be explainable by other factors such as motion blur as well,” they said, concluding that the media “is likely manipulated” through AI.
Video 4
Audio engineer’s analysis
For comparison purposes, Sample A is the claim and Sample B is Kumar’s video from 5 January 2017. Azher explained that when the acoustic behaviours are examined, “the evidence points to clearly different recording conditions and processing for the two files”.
“The bass-heavy spectrum, higher absolute noise floor, higher spectral flux, much shorter decay time, and compressed dynamic range all indicate that Sample A does not share the same ambience character or production chain as the original.
“Conversely, Sample B is brighter, has greater dynamic range, and features a longer natural decay. It should be noted that the spoken content in the two clips is entirely different, confirming they are not the same source recording,” he said.
To support his findings, Azher provided the following observations:
- Voice fingerprint: When the underlying mathematical structure of both audio clips is analysed, the overall MFCC profile shows a clear divergence in the spectral envelope and timbre between the two recordings. Furthermore, Sample A displays a higher mean in spectral flux, indicating a higher rate of spectral change over time, compared to Sample B.
- Dynamic range and compression: Sample A has a lower dynamic range with a crest factor of approximately 4.0, making it significantly more heavily compressed. In contrast, Sample B has a crest factor of roughly 6.7, showing a greater dynamic range.
- Mouth clicks and processing artefacts: Sample B exhibits higher overall activity in its zero crossing rate than Sample A. Sample B also has a higher high-frequency energy ratio and contains slightly more frequent high-frequency transient bursts, around 1.29 per second, compared to Sample A’s 1.11 per second. This is consistent with Sample B being a more open and less processed recording.
- Fake background static: The absolute noise floor level is higher in Sample A. The spectral shape of the quiet parts differs markedly, with Sample A being almost pure low-frequency dominant — about 96% low energy — while Sample B is mid-frequency dominant at about 63% mid energy. This indicates different ambient conditions or processing rather than sharing identical embedded room tone.
- Unnatural pitch: Sample A exhibits a darker, more bass-heavy spectrum, possessing a spectral centroid of approximately 1,358 Hz. By comparison, Sample B is markedly brighter with an overall brighter balance, possessing a spectral centroid of roughly 2,545 Hz.
Deepfake detectors’ results
Soch Fact Check used two phrases from the audio component in Hiya Deepfake Voice Detector: “How much love do you have for Imran Khan? I have a lot of love…” and “Akshay Kumar love Imran Khan. Now fill up the comment.” The tool said both samples were “likely authentic”, with scores of 99 and 98 out of 100.
According to Hive Detect’s results, the sample is 3.2% likely to be AI-generated video, 0% likely to have AI-generated speech, and 0.5% likely to be deepfake.
DeepFake-O-Meter’s calculations were “inconclusive”, with “very low” confidence. The six detectors we used found probabilities of 100%, 29.5%, 1.7%, 40.1%, 66.7%, and 98%.
On the other hand, NVIDIA Synthetic Video Detector said the clip was “synthetic”, with a score of 80.0%, which is “greater than threshold (30%), meaning it’s likely synthetic or AI-generated”.
Moreover, GODDS — which employed 22 deepfake detection algorithms for the visual content and 63 for the audio component, while two trained analysts also examined the clip — also found the sample to be AI-manipulated.
All predictive models for the visual and audio content said the video “is likely to be fake”:
- The video is likely to be fake with a probability above 0.5, according to eight of the 22 predictive models; it is likely to be fake with a probability below 0.5, according to the 14 other predictive models.
- The audio is likely to be fake with a probability above 0.5, according to 61 of the 70 predictive models; it is likely to be fake with a probability below 0.5, according to the two remaining predictive models.
According to GODDS’ human analysts, the video contains “several indicators” that show it may be artificially manipulated. For example:
- 0:00-0:01 The eyebrows appear as black blurs.
- There is an unidentifiable man in the bottom region of the video.
- The lip movements do not align with the sounds.
“The artefacts are largely obscured by poor video quality. Hence, it is difficult to discern possible artefact manipulation from the video’s heavy compression. These indicators can plausibly be explainable by other factors such as motion blur as well,” the analysts said, adding that they “believe this media is likely manipulated” via AI.
Despite the deepfake detectors’ diverse results, Soch Fact Check concludes that all four videos have been manipulated using AI tools as we were able to trace the source footage and the audio engineer’s analyses prove they were doctored.
Virality
The AI-manipulated videos of Akshay Kumar have gained hundreds of thousands of views on Facebook.
One of the clips was also posted on TikTok in May 2023.
Conclusion: All of the videos are AI-manipulated; none of the original clips show Akshay Kumar speaking in favour of Khan.
Background image in cover photo: Soch Videos graphic
To appeal against our fact-check, please send an email to appeals@sochfactcheck.com