New needs for VoiceAI preprocessing by old tools
Stuart Naylor <[email protected]> Thu, 20 Aug 2020 07:07:47 +0000
| Newsgroups | gmane.comp.audio.compression.speex.devel |
|---|---|
| Message-ID | <DB7P191MB03322B78E8A8A59D3679EE8FA85A0@DB7P191MB0332.EURP191.PROD.OUTLOOK.COM> |
--===============3788620128487000883==
Content-Language: en-GB
Content-Type: multipart/alternative;
boundary="_000_DB7P191MB03322B78E8A8A59D3679EE8FA85A0DB7P191MB0332EURP_"
--_000_DB7P191MB03322B78E8A8A59D3679EE8FA85A0DB7P191MB0332EURP_
Content-Type: text/plain; charset="Windows-1252"
Content-Transfer-Encoding: quoted-printable
Hi,
Xiph has done some really good work but since Speex things have seemed quit=
e quiet.
Devs to hobbyists have a huge battle with VoiceAI audio processing and even=
when available extremely similar routines are separated across libs so tha=
t much process duplication occurs and increases load.
The Alsa-plugins of Speex echo just don=92t seem to work whilst repos using=
the speex libs do.
https://github.com/voice-engine/ec
I actually use the speex AGC Alsa-plugin as its a great tool and then run E=
cho via the above.
Coupled to WebRTC_VAD and a python based MFCC routine.
Beamforming and DOA I have given up hope for as I am running so many extrem=
ely similar FFT routines concurrently in different process space that my lo=
ad usage makes me believe there isn=92t a chance.
All of them have extremely similar frame based spectra analysis collection =
that collate and process over different windows for different process but m=
uch of the input process are almost identical heavy load FFT process that f=
or some reason we duplicate because the Linux landscape is one of scattered=
audio processing libs.
Linux currently has much work in terms of open source ASR, TTS, TensorFlow =
frameworks but what we have available in terms of audio pre-processing are =
scattered individual segments of the audio chain and create a seriously ine=
fficient chain.
These are old tools but the whole chain to MFCC output just doesn=92t exist=
.
Stuart
Sent from Mail<https://go.microsoft.com/fwlink/?LinkId=3D550986> for Window=
s 10
--_000_DB7P191MB03322B78E8A8A59D3679EE8FA85A0DB7P191MB0332EURP_
Content-Type: text/html; charset="Windows-1252"
Content-Transfer-Encoding: quoted-printable
<html xmlns:o=3D"urn:schemas-microsoft-com:office:office" xmlns:w=3D"urn:sc=
hemas-microsoft-com:office:word" xmlns:m=3D"http://schemas.microsoft.com/of=
fice/2004/12/omml" xmlns=3D"http://www.w3.org/TR/REC-html40">
<head>
<meta http-equiv=3D"Content-Type" content=3D"text/html; charset=3DWindows-1=
252">
<meta name=3D"Generator" content=3D"Microsoft Word 15 (filtered medium)">
<style><!--
/* Font Definitions */
@font-face
{font-family:"Cambria Math";
panose-1:2 4 5 3 5 4 6 3 2 4;}
@font-face
{font-family:Calibri;
panose-1:2 15 5 2 2 2 4 3 2 4;}
/* Style Definitions */
p.MsoNormal, li.MsoNormal, div.MsoNormal
{margin:0cm;
font-size:11.0pt;
font-family:"Calibri",sans-serif;}
a:link, span.MsoHyperlink
{mso-style-priority:99;
color:blue;
text-decoration:underline;}
.MsoChpDefault
{mso-style-type:export-only;}
@page WordSection1
{size:612.0pt 792.0pt;
margin:72.0pt 72.0pt 72.0pt 72.0pt;}
div.WordSection1
{page:WordSection1;}
--></style>
</head>
<body lang=3D"EN-GB" link=3D"blue" vlink=3D"#954F72">
<div class=3D"WordSection1">
<p class=3D"MsoNormal">Hi,</p>
<p class=3D"MsoNormal"><br>
Xiph has done some really good work but since Speex things have seemed quit=
e quiet.<br>
Devs to hobbyists have a huge battle with VoiceAI audio processing and even=
when available extremely similar routines are separated across libs so tha=
t much process duplication occurs and increases load.<br>
<br>
The Alsa-plugins of Speex echo just don=92t seem to work whilst repos using=
the speex libs do.<br>
<a href=3D"https://github.com/voice-engine/ec">https://github.com/voice-eng=
ine/ec</a><br>
<br>
I actually use the speex AGC Alsa-plugin as its a great tool and then run E=
cho via the above.<br>
Coupled to WebRTC_VAD and a python based MFCC routine.<br>
Beamforming and DOA I have given up hope for as I am running so many extrem=
ely similar FFT routines concurrently in different process space that my lo=
ad usage makes me believe there isn=92t a chance.<br>
<br>
All of them have extremely similar frame based spectra analysis collection =
that collate and process over different windows for different process but m=
uch of the input process are almost identical heavy load FFT process that f=
or some reason we duplicate because
the Linux landscape is one of scattered audio processing libs.<br>
<br>
Linux currently has much work in terms of open source ASR, TTS, TensorFlow =
frameworks but what we have available in terms of audio pre-processing are =
scattered individual segments of the audio chain and create a seriously ine=
fficient chain.<br>
These are old tools but the whole chain to MFCC output just doesn=92t exist=
.<br>
<br>
Stuart<br>
<br>
</p>
<p class=3D"MsoNormal">Sent from <a href=3D"https://go.microsoft.com/fwlink=
/?LinkId=3D550986">
Mail</a> for Windows 10</p>
<p class=3D"MsoNormal"><o:p> </o:p></p>
</div>
</body>
</html>
--_000_DB7P191MB03322B78E8A8A59D3679EE8FA85A0DB7P191MB0332EURP_--
--===============3788620128487000883==
Content-Type: text/plain; charset="utf-8"
MIME-Version: 1.0
Content-Transfer-Encoding: base64
Content-Disposition: inline
X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KU3BlZXgtZGV2
IG1haWxpbmcgbGlzdApTcGVleC1kZXZAeGlwaC5vcmcKaHR0cDovL2xpc3RzLnhpcGgub3JnL21h
aWxtYW4vbGlzdGluZm8vc3BlZXgtZGV2Cg==
--===============3788620128487000883==--