Re: [OSCOM] Open source CMS / documents are content/
"Content-wire Research" <[email protected]>
| Newsgroups | gmane.comp.cms.oscom |
|---|---|
| Message-ID | <008501c52af8$710f20c0$4bfefea9@paola> |
> Hi Paola, Hi Seth! > > My opinion is that content has the hope of being at least semi-structured. More than hope, its a fact. Systems have been producing structured content form WYSIWYG editors for years (well designed systems that is) That is, it can be stored in a structure that isolates particular attributes (for example title, summary, body...) so that > they can be rendered in different ways. yes, that was exactly my point, I think its the main concept behind single sourcing >To me, the word "Document" > describes a file that includes, the content, >format, and the > instructions to render >that format (Microsoft Word, PDF, >MOV, etc.) and the best you can do for managing structured attributes is > >metadata which may replicate some of the text/information that appears >within the document itself. Thats one thing, metadata will tell me what is in that file. What about file conversion? What about if I want to extract the content of one file and render it in another format, which is a perfectly legitimate requirement for an end user? I mean sometimes I have used systems that have cost so much time and money to develop, and they are not capable of performing anything outside the most trivial functions not because of the limitations of the software, but because whoever buit the system was not capable of thinking, did not know anything about technology nor its potential, and did not even understand anything about the professional tasks that the system was supposed to perform Basically, a system is only as good as its architect, and there many people out ther who really do not have a clue. Have you ever used a system and wondered ' who on earth designed this without thinking of this possible feature?' I have, so many times > > So a content creator can either author content into a Content Management > System (which stores the information in a structured repository) or a > document that then can be put into a document system which stores the > document. Well, a content creator working efficiently authors content in a single source format that can be rendered in a variety of ways, as you point out above. But unfortunately some time we must work with 'given' preconditions Assuming that our requirement, like Srinivast, is only 'handle documents' for now. How do we go about that? It depends how good the developer can think We should look at the entire environment, does the organisation have other systems in place? Where and how are the file stored? What are the operational functions that system should perform? What changes are likely to take place in the sofware industry and to the organisation in the foreseeable future? There s a lot going on I agree with you not everybody is interested in building clever and futureproof software And thats the main problem that I see in our industry today > > There are various strategies for taking an asset that has been authored as > a document (where the author is playing the roles of designer and writer) > and programmatically parse through it to extract the content. OK > However, those strategies are prone to errors because the author might > fiddle with their authoring program to achieve a visual effect but > inadvertantly confound the parsing engine. In that case the parser needs to be tweaked to handle exceptions, nothing to do with the overall system architecture A typical example of this is > a MS Word user overrides the font of an H1 style to achieve the look of > a paragraph body. In this case the parsing engine might think that > > whole paragraph was a section title. A system, generally speaking, should be built to handle all generic instances, and then configured to handle exceptions, and not viceversa. These days, in order to maximise results, I recommend a comprehensive analysis of the situation before investing any resource. You may regret not having thought ahead >>>>>>Hope this helps the dialogue Its good to talk Paola DM > > > > Content-wire Research wrote: >> Thanks for the opportunity to discuss CMS and DM, as it has been bugging >> me lately >> >> I am trying to establish the relationship between various classes of >> systems which have >> similar functionalities, with a view to come up with a better defined >> ontology for our domain >> This mean, lets try to understand what this 'overlap' consists of >> >> It all starts with having solid, agreed upon definitions, which we still >> working on >> >> Here is my reasoning, recently circulated to CMPros which has not yielded >> any response/comments >> >> Proposed definitions for >> ------------------------------ >> >>> Content = Organisational Information and Data collected from internal >>> and external sources, stored in any format (text, graphical multimedia, >>> digital and non digital). Content can be structured (documents, >>> records) and unstructured (discussion groups, bulletin boards) >> >> >>> Document? = Content formatted in a specific file type (.doc versus >>> .pdf) >> >> >>> File type = File extension determined by the software package it is >>> produced with >> >> >> >> 4) if the defintions above are true, then >> documents are content >> then content management includes document management >> although document management may require file type specific >> functionalities >> >> Please agree or disagree/comment >> >> >> The above discussion applies to the post below by srinivas/oliver becauz >> when you think of application development of this sort, you've gotta >> think lateral >> I say >> >> Unless you have completely ruled out that your documents of today (words >> and pdfs) will not >> be your content of tomorrow (html/xml) and viceversa, then you'd be >> wasting your time >> thinking of them separately and not as a single application that can >> handle format independent data >> >> Format independent means that anything that the data can be outputted to >> more than one desired format >> >> >> Dont you think? >> >> Paola Di Maio >> >> >> >> >> ----- Original Message ----- From: "Oliver Crow" <[email protected]> >> To: "srinivas mohan" <[email protected]> >> Cc: <[email protected]> >> Sent: Wednesday, March 16, 2005 6:35 PM >> Subject: Re: [OSCOM] Open source CMS >> >> >>> >>> >>> On Tue, 16 Mar 2005, srinivas mohan wrote: >>> >>>> I am very new to CMS concepts.. >>> >>> ... >>> >>>> The application i am intrested to build is to manage word and pdf >>>> documents with support to workflow mgmt and content distribution. >>> >>> >>> >>> This sounds to me like a "document management system" more than a >>> "content management system". I guess there's a lot of overlap between >>> the two. >>> >>> My understanding is that document management systems are primarily >>> concerned with the production and management of documents. They provide >>> workflow tools, archival, searching, and are usually agnostic about the >>> format of the stored content. >>> >>> CMSs are primarily concerned with content publishing. They provide many >>> features similar to document management, and also provide content >>> editing tools. They typically require that the content be stored in a >>> format editable and manipulable by the CMS, which usually means not Word >>> or PDF. >>> >>> If your application is more about production and management of documents >>> than it is about publishing documents, perhaps the assumptions and >>> feature sets found in document management systems would be a better >>> match. >>> >>> >>> Oliver >>> >>> _______________________________________________ >>> General mailing list >>> [email protected] >>> http://oscom.org/cgi-bin/mailman/listinfo/general >>> >>> >> >> >> _______________________________________________ >> General mailing list >> [email protected] >> http://oscom.org/cgi-bin/mailman/listinfo/general >> > > -- > Seth Gottlieb, Optaros, Inc. > 155 Second Street, Cambridge, MA 02142 > 617.225.2455(v), 617.852.2956(m) > >