Skip to content
View pujariaditya's full-sized avatar

Block or report pujariaditya

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
pujariaditya/README.md

whiteboard banner

typing

badges

about me

i work on trustworthy audio — watermarking voices, tracing synthetic speech, and taking models apart from the inside (TTS internals, neural codecs). if an audio metric can be broken, i want to be the one who breaks it... then fixes it.

divider

research threads

AudioAuth — dual-watermarking framework: frequency-partitioned model + data watermarks for audio integrity and source attribution (IEEE TBIOM 2026)

WaveVerify — FiLM-generator + MoE-detector watermarking; zero BER under common distortions; beats AudioSeal and WavMark (IJCB 2025)

sourcetrace — codec-residual open-set source tracing of audio deepfakes (EnCodec residual + frozen WavLM-Large; FPR95 1.14% vs published 3.36%)

spanmark — two-route segment localization of partially spoofed speech; cross-corpus SOTA on LlamaPartialSpoof (29.0 EER vs 35.5 baseline)

wavepainter — multimodal LLM-guided diffusion for text-based speech editing; substitution WER 3.48 vs prior best 4.41

cond-ID — speaker unlearning in zero-shot TTS via conditioning-space identity redirection (XTTS-v2, Tortoise-TTS, IndexTTS-1.5)

HiggsAudiov2TokenizerUnofficial — full PyTorch training pipeline for the Higgs Audio V2 tokenizer (HuBERT semantics + DAC + 8-layer RVQ, 960x downsampling)

divider

the toolbox

toolbox sticky notes

divider

by the numbers

GitHub streak stats

pinned to the board — updates live

divider

footer

Pinned Loading

  1. wavepainter wavepainter Public

    wavepainter: multimodal LLM-guided diffusion for text-based speech editing — edit a word in the transcript and only that span of the spectrogram is re-predicted

    Python 1

  2. CoRA CoRA Public

    Codec-residual open-set source tracing for audio deepfakes

    Python 1

  3. cond-id-tts-unlearning cond-id-tts-unlearning Public

    cond-ID: conditioning-space identity redirection for speaker unlearning in zero-shot TTS — one speaker-conditioning edit across XTTS-v2, Tortoise-TTS and IndexTTS-1.5

    Python 1

  4. HiggsAudiov2TokenizerUnofficial HiggsAudiov2TokenizerUnofficial Public

    Unofficial PyTorch implementation of Higgs Audio V2 Tokenizer with HuBERT semantic features. Complete training pipeline for semantic-acoustic audio tokenization with 960x downsampling and 8-layer RVQ.

    Python 6 2

  5. AudioAuth AudioAuth Public

    A dual-watermarking framework for robust audio integrity verification and source attribution.

    Python 1