First of all, thank you for creating and maintaining Silero VAD and making it freely available to the community.
I tested Silero VAD on three Russian interview recordings, totaling about 459 seconds, and compared the results with FireRedVAD.
FireRedVAD detected 18 silence boundaries inside regions that Silero returned as continuous speech. At those boundaries, FireRedVAD remained below its threshold for 310–600 ms, exceeding the configured minimum silence duration of 250 ms.
Over the same intervals, Silero remained below its default exit threshold of 0.35 for only 0–224 ms, and in several cases did not fall below 0.35 at all.
As a result, Silero sometimes merged multiple speech turns into continuous segments lasting 64–71 seconds despite audible pauses longer than 250 ms.
This concerns silence-boundary detection, not speaker diarization.
Could you please clarify whether this is expected behavior, a known model limitation, or a possible issue that should be investigated?
I have attached the WAV files, JSON reports, and configurations needed for analysis.
Thank you for your time and for your work on this project.
audio.zip
vad_reports.zip
First of all, thank you for creating and maintaining Silero VAD and making it freely available to the community.
I tested Silero VAD on three Russian interview recordings, totaling about 459 seconds, and compared the results with FireRedVAD.
FireRedVAD detected 18 silence boundaries inside regions that Silero returned as continuous speech. At those boundaries, FireRedVAD remained below its threshold for 310–600 ms, exceeding the configured minimum silence duration of 250 ms.
Over the same intervals, Silero remained below its default exit threshold of 0.35 for only 0–224 ms, and in several cases did not fall below 0.35 at all.
As a result, Silero sometimes merged multiple speech turns into continuous segments lasting 64–71 seconds despite audible pauses longer than 250 ms.
This concerns silence-boundary detection, not speaker diarization.
Could you please clarify whether this is expected behavior, a known model limitation, or a possible issue that should be investigated?
I have attached the WAV files, JSON reports, and configurations needed for analysis.
Thank you for your time and for your work on this project.
audio.zip
vad_reports.zip