Re: Comments on "A File Aggregation Scheme for FLUTE": draft-neumann-rmt-flute-file-aggregation-01.txt
"Christoph Neumann" <[email protected]>
| Newsgroups | gmane.ietf.rmt |
|---|---|
| Organization | INRIA |
| Message-ID | <op.su6t7hi1wcfw6x@goofy> |
Hi Mike, hi all, excellent comments Mike, great to see that you see the value and relevance of his work. Pretty much all of your comments were spot on. See comments in-line. The issues raised will be included in future versions of the draft. On Wed, 20 Jul 2005 00:46:01 +0200, Michael Luby <[email protected]> wrote: > Comments on “A File Aggregation Scheme for FLUTE”: > draft-neumann-rmt-flute-file-aggregation-01.txt > > Section 1.1.2, in first paragraph after list of points, two sentences at > the > end say “It is also important to note that the physical object > aggregation > scheme is NOT RECOMMENDED for “large” files, i.e., files of a size > exceeding > a given threshold that depends on the FEC instance used. This issue is > also > discussed in Section 4.6” The first sentence is vague and I don’t think > the > rationale is always correct, so I suggest deleting these two sentences. > It > already says that recommendations on file sizes and when to aggregate > are to > be made in Section 4.6 in the previous sentences, and this is more than > enough for this introductory text. *** The 2 sentences are maybe are bit to vague. However, this is the intro and it ought to state the applicability of this document - so it's important to know that this document is most relevant if you're not talking about exclusively large files. So we propose keeping the applicability but in a bit shorter form. We propose the following rewording: "The physical aggegration file aggregation scheme is mainly applicable for small files." > > Sections 4.2 and 4.4.1, It is not clear from this text, but it would be a > big advantage if it were possible for each file within the aggregate > file to > start on a symbol boundary instead of having one file end and another > start > within the middle of a source symbol. (And this might even be > RECOMMENDED.) > This would mean each source packet would contain information about only > one > file, and not a mix of files, whereas if one file ended in the middle of > a > source symbol and the next file started immediately then the packet > containing this source symbol would contain a mix of information about > these > two files. *** Excellent idea. We have to look on how we can signal this to the receivers. An optional attribute SymbolAligned="true/false" for the element of the aggregated object should do it... > Thus, it would be good when specifying the offset of a file into > the aggregate file that it could be on a symbol boundary (and it might > even > be good to mandate this, but this is a different discussion). It is not > clear that this is possible with the current spec, since I didn’t see any > discussion of what to do if there are bytes between the end of one file > and > the start of the next. If this suggestion is taken, then one would have > to > explain that any bytes between the end of one file and the beginning of > the > next file in the aggregate file MUST be logically considered to be zero > when > generating repair symbols from the aggregate file. Note that it should > also > specify that these padding bytes in the last source symbol of a file > don’t > need to be sent in the source packets (and thus the payload of the last > source packet sent for a file might be truncated and might not be a > multiple > of the symbol length), but are only required to generate repair > symbols/packets. *** All this sounds good to me, we have to look at the mechanisms and change needed, and include this in the next version of the draft. > > Section 4.6.1, point 1 in the list, it says “The aggregated object size > SHOULD NOT exceed the FEC maximum source block size. This prevents from > files to straddle several blocks.” This is not necessarily sound > advice, > as there are distinct advantages to allowing the file transport to send > one > aggregate file spanning multiple source blocks versus multiple aggregate > files spanning one source block each. One advantage is that since the > packets for the source blocks are interleaved, when there are burst > losses > the losses hit all source blocks equally and this spreads out the losses > nicely for the one aggregate file, whereas interleaving would have to be > done at the application layer to have this same advantage when sending > multiple aggregate files of one source block each. If this is a SHOULD > NOT > then this scalability advantage is lost, and it puts the onus of trying > to > secure this advantage on application. Furthermore, there is more > application > bookkeeping overhead with managing multiple aggregated files versus a > single > aggregated file. Instead of the current SHOULD NOT, a more informative > discussion of the advantages and disadvantages of each approach without > any > normative advice seems like it would be more appropriate. > > Section 4.6.1, point 2 in the list. This point is essentially meant to > be > the same as point 1, just said differently. Thus, advice is to delete it > and provide informational only advice as described above. If this point > is > kept, then at the very least it should be correct to say what is intended > (right now it says that the aggregated object should be as close as > possible > to the FEC maximum source block size, which is silly if the total > aggregate > of all the files that are meant to be sent is much smaller than this > size, > and the current wording may implicitly suggest inefficiencies, such as > padding the aggregate file size out with zeroes to make it large enough, > or > waiting some indeterminate amount of time until the aggregate size of > files > to be sent is close to the maximum source block size.) > > Section 4.6.2, Point 1. Same comments as above, i.e., the advice in the > first sentence is not always correct and should not be stated this > strongly. > The second sentence also seems overly restrictive, as any reasonable > blocking algorithm will automatically partition the aggregate file into > source blocks that are as long as possible and as equal length as > possible. > Not sure why this is any different than an algorithm for partitioning a > single file into source blocks, where the same strong advice is not given > (because it is not needed, the partitioning algorithms automatically do > the > right thing). *** For the entire section 4.6, you are certainly right that it lacks a bit of work here, and is a bit to restrictive. A more informative discussion on the objects sizes, the advantages and disadvantages is necessary. We think that a detailled discussion on this should be done in a seperate draft... The intention is to first test it out, and look at the results (as stated in the editorial note). This seems apropriate if we wan't to go for an experimental RFC. > > Section 4.7, Point 1. This point suggests that there are exactly two > ways > to update a file. This is not necessarily correct, and this language > should > be weakened to say that there are at least two ways. Also, in the > description of the first method, it says that a brand new aggregated > object > is created and replaces the previous aggregated object whose transmission > MUST stop. The “MUST” seems inappropriate, as this is up to the > application > and its update rules (even “SHOULD” seems to strong here). This should > be > more descriptive than normative text. You are right. The wording in the draft is maybe a bit to restrictive. We should state: "If an update is required at least two possibilities exists. Other mechanism may be used, but are out of scope of this specification." > Section 4.7, Sentence after Point 1. Not clear why the TOI MUST start > with > 1 and be incremented by exactly 1, and there might be reasons why this is > undesirable. Stating that it MUST start with a value that is at least > one > and increments by at least one seems more flexible and acceptable. Same > comment applies in Section 5.1 in the last paragraph, where it should > also > state that handling wraparound of the TOI space is out of the scope of > this > document. Ok for the more flexible approach stating that the TOI should start with a value that is at least one. However we can see no reason why we should not increment the TOI by exactly 1. So we propose to put: "To that purpose the TOI assigned by the sender to each object MUST start with at least 1 and be incremented by one for each new object. This ensures that the receiver can unambiguiously determine which instance of a certain file URI is not obsolete (the one with the logically highest toi). This has no implication on sending or receiving order, only on allocation." cheers, Christoph, Vincent, Rod