RE: Should FLUTE File Aggregation I-D be a working group item?
"Michael Luby" <[email protected]>
| Newsgroups | gmane.ietf.rmt |
|---|---|
| Message-ID | <200512091115.jB9BFHsB021777@mail> |
A couple of comments below. -----Original Message----- From: [email protected] [mailto:[email protected]] Sent: Friday, December 09, 2005 2:23 AM To: [email protected]; [email protected] Cc: [email protected]; [email protected]; [email protected]; [email protected] Subject: RE: [Rmt] Should FLUTE File Aggregation I-D be a working group item? Sending a 99% identical large multipart object for a 1% is indeed a very bad efficiency. Not that you wouldn't do it anyway if you were delivering in a file carousel mode. *** I didn't mean to suggest sending a 99% file for updating 1%, sorry if I gave that impression. However, you are right that this can happen, if you are continually refreshing a set of files, but you would hope that you would set the refresh rate in such a way that you aren't that inefficient. The topic is sufficiently separate to merit a separate I-D so maybe we don't need any thorough description of it under the AO I-D's physical aggregation text. So I think Mike's suggestion is the way to go. *** I'm sorry that I missed that it had already been decided to roll the incremental update feature into the main AO document. Given that this has already been agreed, I think that is the way to go. I'll leave it to the coauthors to figure out how to roll this feature in and then comment on it afterwards. Cheers, Rod. >-----Original Message----- >From: ext Michael Luby [mailto:[email protected]] >Sent: 08 December, 2005 17:44 >To: Walsh Rod (Nokia-NRC/Tampere); [email protected] >Cc: [email protected]; [email protected]; [email protected]; >[email protected]; [email protected] >Subject: RE: [Rmt] Should FLUTE File Aggregation I-D be a >working group item? > >Hi Rod, >There are a bunch of issues with what you suggest, and perhaps >the most important is that this is a very inefficient way to >do incremental file updates. What I suggest as the way >forward is for somebody (you can guess >who) submit an individual submission that makes a proposal on >how to do incremental updates, and then we can discuss in the >working group and if there is acceptance this can be rolled >into a file aggregate working group item once that is >accepted. I'm not sure it is useful to have email discussions >on this as it seems more efficient to have a concrete proposal >to discuss. >Mike > >-----Original Message----- >From: [email protected] [mailto:[email protected]] On >Behalf Of [email protected] >Sent: Thursday, December 08, 2005 7:10 AM >To: [email protected] >Cc: [email protected]; [email protected]; [email protected]; >[email protected] >Subject: RE: [Rmt] Should FLUTE File Aggregation I-D be a >working group item? > >Hi Christoph, Mark, et al. > >Just on the "incremental update" issue... > >The obvious method is to ensure that some of the source blocks >are identical between AO x and AO x+1, and with the same >FEC-OTI it's trivial to use source blocks of AO x with AO x+1. >It's important to specify a single simple way for the client >to recognise this feature. >Since the client gets the AO structure anyway as byte ranges >(it must do for partial recovery), then it can easily >calculate which blocks are the same for the same AO (i.e. AO >using the same fileURI). For different source blocks it just >receives symbols for the new one (AO x aren't of much use to >those source blocks in this scheme). The question is whether >to make this FEC code specific. > >For some FEC codes (e.g. if you take the whole AO as one >massive source block), you'll need to rely on an understanding >of "where the encoded symbols came from" (e.g. the PCM). > >Probably the easiest way to deal with this in the I-D is to >explain the simple way to do this for identical blocks, and >then explain that mapping redundant symbols back to source >symbols could be used for different blocks with some common >(file) data, but leaving the exact details of mapping back >something like a PCM out of scope. Then all angles are covered >without any need for rocket science. > >Sound good? > >Cheers, Rod. > >Ps Since I'm not a FEC expert, please be kind about any >glaringly stupid suggestions on FEC :) > > > >________________________________ > > From: [email protected] >[mailto:[email protected]] On Behalf Of ext >Christoph Neumann > Sent: 06 December, 2005 00:32 > To: Walsh Rod (Nokia-NRC/Tampere) > Cc: [email protected]; [email protected]; >[email protected]; [email protected] > Subject: Re: [Rmt] Should FLUTE File Aggregation I-D be >a working group item? > > > Hi all, > > sorry for my late reply to all your comments. I think >quite interesting issues have been raised. > I will only talk about physical aggregation, since the >logical aggregation scheme seems quite obvious to everyone. > > First, what is clear to all of us, is that the object >aggregation scheme should be more than just a gzip or tar of >several files. > > Therefore, we need to clearly identify the advantages >we want and can provide. From the previous mails I identfied >the following >features: > I. Partial reception: Extraction of files from >partially recovered AO > > II. Improved transmission performances (eliminating >coupon collector + better FEC performances) > > III. Provide FDT information (MIME-type, URI etc.) for >each of the files that have been aggregated > > IV. Backward compatibility: > This feature can be interpreted in two ways: > 1) Receivers which do not support the aggregation >scheme, should be able to interpret at least the source >symbols sent for the files being aggregated. > 2) Receivers which do not support the aggregation >scheme, should be able to interpret all source and parity >symbols of the aggregated object. I.e the receivers can >reconstruct the entire aggregated object, and then decompose >it into the files being aggregated. > > V. Incremental update: It should be possible to reuse >previously received symbols of an old aggregated object or of >old files, when these files are reused in a new aggregated object. > > VI. Single file transmission: It should be possible to send the >files being aggregated individually. > > > The "gzip approach", only supports features II. and IV.2). > > The draft presents a solution that supports feature I. II. III. >IV.2) and nothing more. > The MIME format was chosen primarily with issue IV.2) in mind. >As mentioned in your comments it is maybe not the best >solution and we may have a look at some alternatives. > > Another proposal that came out of the discussions looks >interesting to me: sending each individual file on a separate >TOI and having an aggregated object that uses another TOI. > The parity symbols produced for the aggregated object >are generated based on all source symbols of the files being >aggregated. >This means that only parity symbols are send for the AO. > With that approach the FDT must list the TOIs of the >files being aggregated and their offsets within the aggregated >object (maybe some additional attributes will also be needed). > As mentioned by Mark, the FDT would be much simpler >compared to the current approach proposed in the draft. >However, we must clearly think about the mechanisms used by >such an approach. > > This new approach should be able to cover the features I. II. >III. IV.1) V. and VI. > The only concern is that receivers which do not support >the aggregation scheme cannot use the parity symbols > of the aggregated object. However I think this is >acceptable, since the sender may produce and send parity >symbols for each individual file, if he suspects having >receivers that do not support the aggregation scheme. > > I think it is worth investigating this proposal, and I >am a volunteer for doing that. As pointed out by Mark, the >resulting draft will be rather different from the one >presently proposed. With regards to the comments in this >thread, I think that the new document should then become a >working group item. > > cheers, > > Christoph > > > On 12/2/05, [email protected] < [email protected] ><mailto:[email protected]> > wrote: > > This is an excellent point. Basically we do not >want to specify > applications in RMT and vigilance is essential. > > (I'll skip over the logical aggregation chatter >as the thread is focused > on physical aggregation). > > FLUTE already allows content encoding for >transport - great idea taken > from HTTP. FDTs provide the necessary >identification for files and the > CENC header extension for FDT Instances. As >such it is possible to > "code" (e.g. zlib/gzip) a single file. For >doing this to a group of > files you have but 2 options: > - gzip each file and transport separately > - gzip together in an application and reveal >only the .zip file & MIME > type to FLUTE (presumably with no content >encoding in this case) > > These 2 options are fine in their respective >use cases; and note I'm > using zlib and gzip only to make the discussion clearer >- there are > several other examples which may be relevant (tarball, >...) > > The problem is the 2 options force you to lose >out on transport > functionality. You either forget FEC >performance (many small files) or > you forget file delivery (i.e. FLUTE becomes a >container shipment > activity with no knowledge of files). > > If you need the atomic file deliver primary >function of FLUTE, as well > as the FEC performance of larger objects then >you'll need a 3rd option > at the transport layer = the physical >aggregation technique. > > A really exciting part is that although >transport grouping binds files > together (in an all-or-nothing sense); with >partial recovery you would > not only be able to serve all-or-nothing applications. >i.e. handling > this as transport really can improve transport >performance in ways that > "above the FEC" application functions would not. > > Cheers, Rod. > > PS if gzip provided the right data about files >in an external/exportable > MIME type and allowed partial recovery, then >I'd go for simple adoption > of gzip, but it doesn't - so I think we can >either specify this simple > object aggregation, or we create gzip2 which >does have the features we > need at FLUTE transport. I don't like the idea >of inventing gzip either, > and I'm not too keen on evangelising the >adoption of re-invented gzip2 > :) > > > > >-----Original Message----- > >From: ext Joerg Ott [mailto: [email protected]] > >Sent: 28 November, 2005 10:09 > >To: Walsh Rod (Nokia-NRC/Tampere) > >Cc: [email protected]; >[email protected]; > > [email protected] ><mailto:[email protected]> ; [email protected] > >Subject: Re: [Rmt] Should FLUTE File >Aggregation I-D be a > >working group item? > > > >Hi Rod, > > > >I think fioe aggregation is basically a useful thing to > >consider as the motivation clearly points out. > In Vancouver, > >I raised the point about re-inventing gzip. >The following > >statement in the motivation section makes me a bit >nervous: > > > > Another motivation for having an object aggregation >scheme compared > > to a basic archive based solution (e.g. tarball), >is that no extra > > transformation (i.e. archive creation or >extraction) is required at > > either the FLUTE sender or receiver. Everything is >managed > > automatically by the transport mechanism according >to transport- > > specific optimizations and can be transparent to >upper applications > > (i.e. built on top of FLUTE), or enhanced by >application hints on > > file relationships, without breaking the basic >semantics of FLUTE > > sessions. > > > >Yes, you can always push functionality down a >layer (which > >still means that you have to perform the >function but you just > >abstract from its details for the application >and limit its > >amount of work and the flexibility in applying >this function). > > But you should carefully ask yourself how far >you want to > >take this. Will we see semantically more complex > >relationships tomorrow? Compression of data? ...? >How do > >you deal with partially received objects? > >How complex does the API to your transport get? > > > >There is a certain number of things that only >the application > >can do right by itself. I would like to see a clear > >borderline drawn of what shall go into the >FLUTE transport and > >what should stay up at the application layer >and then we > >should make sure that the extensions we define >now are going > >to be orthogonal to other features we may want >to add to the > >transport later and that they are easily usable by the > >application for what it may want to do. > > > >(Remember that we are not building "systems" >in the Transport > >Area and that we do not necessarily have to >build application > >libraries either. So, how much is needed for >interoperability > >in FLUTE?) > > > >Cheers, > >Joerg > > > >> Hi Mark et al. > >> > >> Thanks for the support > >> > >> <I'll take time to consider the AO semantic >suggestions - an > >email for > >> another day> > >> > >> > >>>Some other minor comments: > >> > >> > >>>- I don't think we need to worry about >backwards compatibility with > >>>FLUTE receivers compliant to the >experimental version which did not > >>>have element extensibility in the FDT XML - >this was a > >mistake and the > >>>'experimental' phase is exactly supposed to >allow us to find and > >>>correct such mistakes. > >> > >> > >> Now the update FLUTE is practically ready >(anticipate one last I-D > >> update), I agree. > >> > >> The AO I-D previously the use of non-TOI=0 >FDTs to help with Exp > >> backwards compatibility. Non-TOI=0 FDTs are >allowed anyway, but we > >> need not explain that this is useful in handling the > >experimental FDT > >> extensibility issue. > >> > >> > >>>- the draft includes some text attaching >semantics to TOI ordering - > >>>something we removed from FLUTE in favour of >FDT Instance IDs. We > >>>should stick with the FLUTE principle that >later version of > >a file is > >>>signalled by using a different TOI in an FDT >with a larger FDT > >>>Instance ID - it doesn't matter whether the >TOI value is greater or > >>>less than the previous one. > >> > >> > >> Note, for FLUTE there is no such semantic. >The client has no > >gauntees > >> about FDT Instance numbers, no garuntees >about TOI numbers. > >> > >> Generally, using TOI and 'FDT Instance >number' values for this has > >> similar issues (assuming wrap-around is >handled smoothly, > >higher/later > >> values represent later versions, gaps in the >range of number are > >> feasible, FDT Instances may never be >received). TOI keeps semantics > >> with the file declaration; FDT requires >inheritence from FDT > >Instance > >> to files described for that Instance. >Implemantation-wise there's > >> little difference between the two: probably >the server will be > >> incrementing FDT Instance numbers and not describing > >out-of-date TOIs > >> on new Instances; probably the server will >use higher TOI values for > >> later versions; the client must store TOI >with fileURI. There is no > >> requierment that the client stores FDT >Instance numbers. > >> > >> Quick FLUTE quote: > >> > >> A receiver of the file delivery session keeps an >FDT database for > >> received file description entries. The receiver >maintains the > >> database, for example, upon reception of FDT >Instances. > >Thus, at any > >> given time the contents of the FDT database >represent the > >receiver's > >> current view of the FDT of the file delivery >session. Since each > >> receiver behaves independently of other receivers, >it > >SHOULD NOT be > >> assumed that the contents of the FDT database are >the same for all > >> the receivers of a given file delivery session. > >> > >> Since FDT database is an abstract concept, the >structure and the > >> maintaining of the FDT database are left to >individual > >> implementations and are thus out of scope of this >specification. > >> ... > >> * The receiver SHOULD NOT use a received FDT >Instance to > >interpret > >> packets received beyond the expiration time of >the FDT > >Instance. > >> > >> i.e. the expiry time of file description >should be kept with other > >> file data, but FDT Instance ID is up to the >implementation. > >> > >> Forcing clients to additional store this >value is not much trouble - > >> but it is _not_ "the FLUTE principle that >later version of a file is > >> signaled by using a different TOI in an FDT >with a larger > >FDT Instance > >> ID". Which is exactly why the AO draft >offers the simplest > >solution to > >> this issue. If we make the assumptions of >"well-behaved > >simpple FLUTE > >> servers" above, there really is nothing >between FDT Instance ID and > >> TOI to be used for this purpose. The reason >I give all this > >background > >> is that this doesn't hold for some obvious >AO usage... > >> > >> For a physically aggregated object (e.g. >very large) with some small > >> updates (e.g. 2 file versions are new from >10,000 files). The server > >> will use TOI (i) for the initial full object files > >descriptions. It is > >> highly likely that the big aggregated object >will continue to be > >> transmitted for some time and with only >minimal updates it > >is probably > >> more bandwidth efficient to send the symbols >for the original object > >> and the few (e.g. 2) new file versions >seperates (more efficient due > >> to, e.g. fountain, symbols of earlier >reception still being > >useful to > >> all receivers regarless of complete, partial >or not-started > >receiption > >> status). So the server sends FDT (ii) with >the new TOIs for > >the files > >> which have been updated. <At this point >using either TOI or FDT > >> Instance ID would work>. However, some objects of FDT > >Insatnce (i) are > >> now changed/removed so that the server need >to describe the slighlty > >> old large aggregated object on TOI (iii). >All recievers > >throw away the > >> new files (even though the TOI is higher) as >the old files are > >> described in a later FDT Instance. > >> > >> FLUTE is specifically designed to allow this >ebhaviour and make no > >> assumption the receivers FDT/file database >state and no > >assumption on > >> reception of individual FDT Instances. > >> > >> Using TOI for this is simple and works. FDT >Instance ID > >fails. Hence, > >> the mention of TOI for this purpose in the I-D. >(Originally it was > >> start with TOI=1 and increment by 1, but >that's overkill as just > >> 'higher value' is needed for subseqent >versions of a file (fileURI) > >> instance) > >> > >> (I've been trying to explain this for a >while - but it's not > >important > >> in 3GPP rel6, and I'm not up-to-date with >DVB or OMA to see if > >> something's fractured there - so maybe this >is the first time I've > >> explained by mail.) > >> > >> > >>>- the draft mentions that one might only >consider file aggregation > >>>when the aggregate is less than a single source block >- I > >think it is > >>>simpler and has no disadvantages if you >decouple the source blocking > >>>(which is a feature of the FEC scheme >chosen) from the file > >>>aggregation. If files belong together in an >aggregate then > >they should > >>>be put together and then if the FEC code >needs to split that > >data into > >>>blocks this is ok and does not result in any worse > >performance so long > >>>as it correctly interleaves the packets from >the different blocks. > >> > >> > >> Exactly. > >> > >> The I-D does not mean to bind AO and FEC blocking. >The > >discussion is for > >> implementation guidence: for files much >greater than the block size > >> there's no FEC benefit in physical >aggregation; for the > >composite files > >> being less than the block size it always >makes sense to aggregate for > >> robustness; inbetween it's a compromise; the >I-D absolutely does not > >> impose any FEC instance retraints and so >it's implementation > >specific to > >> optomise complexity and performance in this aspect. >The guidence is > >> valuable. If we inadvertantly made normative >specification > >on this, it > >> was an error to be corrected. > >> > >> Also, the logical grouping between objects >can be used to get the > >> "belongs together" semantics for multiple AOs - e.g. >where partial > >> recovery is undesirable, or different levels >of robustness are > >> desirable). > >> > >> Cheers, Rod. > >> > >> > >> > >> > >> > >> > >> > >> > >> > >> > >> _______________________________________________ > >> Rmt mailing list > >> [email protected] > >> https://www1.ietf.org/mailman/listinfo/rmt > >> > > > > > > _______________________________________________ > Rmt mailing list > [email protected] > https://www1.ietf.org/mailman/listinfo/rmt > > > > > > >_______________________________________________ >Rmt mailing list >[email protected] >https://www1.ietf.org/mailman/listinfo/rmt > >