Re: Feature removal proposal

Soren H <[email protected]> Sun, 18 Apr 2004 14:44:02 +1000
Newsgroups gmane.comp.audio.speak-freely.general
Message-ID <[email protected]>
> Jeff MacKinnon wrote:
> > I would like to again reiterate my strongest recommendation (which I have
> > brought up in prior posts) to retain a feature that was introduced in 7.6,
> > and which you haven't specifically mentioned in your request for feature
> > removal feedback.  That is, the feature that John added which he called
> > "Adaptive Output (Audio) Rate Adjustment" which addresses the long
> > standing and much reported "increasing audio delay" problem.
> >
> > I bring this up because some previous posts have tended to undervalue its
> > importance, not realizing that many users over a long period of time
> > complained about this problem.  More users than one might think prefer to
> > communicate in full-duplex constant transmit mode, rather than utilize the
> > push-to-talk feature.  It is much more natural and "telephone like."  I've
> > encountered users who refuse to operate any other way.  It would be a real
> > shame to lose this much awaited feature after having finally obtained it
> > in 7.6.

There are a few complications with rate adjustment feature.  As it
operated in 7.6 it was partly a compensation mechanism to cover up
other deeper packet-handling and timing issues.  As such, some see it
as a substitute for packet timestamps and even sequence numbers.
That's only true if you are wanting to avoid the delays of a couple of
seconds or more from mismatches and delays remaining.  For the next
release, I, at least, want to see the delays really minimised - and
well beyond what the 7.6 scheme can allow.  For a natural spontaneous
conversation, you really want less than 1 second round-trip, and that
means 1/2 second each direction.  Using the 7.6 rate adjustment, you
can expect at least 1/2 second just in the audio buffer delay, let
alone the network.

My early packet schemes worked by feeding the audio output only when
needed, allowing the most time for packets to arrive before bridging
techniques would be applied for missing packets.  That meant that the
rate adjustment would not work very effectively when there was
significant packet loss, but on the other hand, it was less necessary
because the packet handling would avoid many of the issues that caused
the delays in the first place.

But then in my efforts to prune the final 200ms or so of delay caused
by windows drivers variability, I altered the scheme to instead
synchronise on the incoming packet stream, relying on the timestamps
for statistics about timing.  This avoided potentially unstable
feedback scenarios when adapting to the delays the wave interface
needs.  So now the audio system is supplied at the rate that the
network stream can reliably deliver continuity at.  I knew that now a
means would be needed to avoid audio queue growth, and found that the
rate adjustment will basically do that job.

It can be improved a bit from where it is now though by making full
use of received packet statistics.  I think that for the adaptive
jitter mode, the rate adjust scheme will be replaced by quiet-packet
dropping when needed, because that will give a faster and more
predictable response.  Something like the quietest 1 in 100 packets
would be candidates to be dropped if there is more audio in the queue
than is needed to sink the incoming stream for the 99th percentile of
packet latency.

The upshot of this is there should be no concern for increasing and
long delays in the next version.  Voice activated mode will still
generally be able to deliver 100ms or so less delay because of the
nature of the windows wave device.  Long-term delays should be
considerably shorter than they are with 7.6, provided both ends have
the new version.  Because of the lack of timestamps in the existing
version, the next release is unlikely to perform any better when
communicating with the existing version at the other end.


> >
> > Also, of lesser importance, I would very much like to see the new (in the
> > Windows version of 7.6) "Systray Icon" feature retained.
> >

Systray icon should live.  I've never actually seen it going myself,
but have noticed it in the code.  It's hidden in #ifdef which accounts
for why its not compiled in by default under the mingw makefiles
(yet).  It seems pretty inoffensive though, and should work when enabled.


On Sat, Apr 17, 2004 at 10:32:25AM +0200, Attila Konietzka wrote:
> I really wonder what's so hard and difficult to utilize voice-activation
> when it's available. I really didn't find it difficult to setup contrary to
> what john walker said in his help-file for speak-freely. It reduces lots of
> bandwidth and the delay-adjustment didn't help that much in my opinion, but
> anyway i think this feature is valuable, i am just not sure, if i was using
> it in a proper way.
> Currently i am using sf 7.6a.
> Regards, Attila.

There's no comfort-noise generation on the receiver, so some people
can be disturbed by the discontinuous nature of the communications.
There are some improvements in my development version that address
the shortcomings of voice activation.

* using speex with denoise enabled, reasonable audio hardware will
give negligible background noise, making comfort noise less an issue.

* I have replaced the VOX with a new one based on the RMS of the
signal over a frame.  Previously unix used the peak level and a
defective timeout, and windows used a mean-of-abs across the frame.
The new one should give much more accurate voice detection.

* In the unix version, there is now a look-ahead on the vox, whereby
one extra packet is sent before the voice activated ones.  This avoids
any cut-off at the start of a talkspurt.

* The timestamps and statistics collection means that initial jitter
delays can be greatly reduced, because you know the pdf for the
arrival time of future packets. 

* With more efficient compression, the headers, including UDP and IP
headers, are a significant network load.  These are reduced with
voice-activation, but are not when just making use of the variable
bit-rate encoding.

* The log-to-ogg-file only works properly with voice-activated, as
otherwise you store a lot of quiet, and the conversation isn't mixed
properly. 

So I think things are turning in favour of voice activation, but there
will also be environments where people don't directly pay for their
network usage level and want constant connection. 


Soren


-------------------------------------------------------
This SF.Net email is sponsored by: IBM Linux Tutorials
Free Linux tutorial presented by Daniel Robbins, President and CEO of
GenToo technologies. Learn everything from fundamentals to system
administration.http://ads.osdn.com/?ad_id=1470&alloc_id=3638&op=click