Re: Re: MF Media Tools project

Kendall Clark <kendall-4GNy1lrxftmrG/[email protected]>
Newsgroups gmane.politics.leftists.monkeyfist
Message-ID <[email protected]>
On Fri, Aug 03, 2001 at 09:46:41AM -0700, kellan wheezed something about:
> Thanks for thinking of me, what an awesome sounding project!

Coordination with IMC is a must, IMO.

> My first thought is this might start to answer the question, "So now
> what?".  The web and activism is still in its infancy in many ways (I had
> a bunch of squatter convincingly explain to me that it takes 10 years for
> a good strategy to mature), and IMC is a product of that infancy.  Its
> still just getting going, but already the limits of its usefulness are
> beginning to come clear.  So I'm thrilled.

Well, frankly, yeah, that's part of the motivation (for me, others I can't
speak for) here; that is, I'm not really crazy about the quality of stuff I
find on IMC. I'm all for open publishing as an idea, it's just that I don't
like wading through junk to find jewels.

But it seems to me we might get better quality analysis if a few things
happened: 1) we had better tools, 2) better control of primary sources, and
3) more consistent, rigorous application of Propaganda Model.

As it happens, I and Bijan and some others like to say we're intellectuals
of the left; if true, we have certain responsibilities to do this kind of
work. 

> Thinking about what would be needed to make this project work, a perhaps
> non-obvious piece is the support for the human editors.  This sounds like
> an editor heavy project, and so ideally you want good tools,
> documentation, and a community for the editors to operate in.  

Absolutely. What I don't talk about so far in that document is people-power;
this isn't some pie-in-sky AI project; we just want to build good tools for
humans to use.

This also implies that this project won't be "open" in the sense that IMC
is. Media analysis isn't particularly hard but it is skilled. One reason to
do this is to train folks in the art and skill of it.

  Not
> necessarily new tools to develop, perhaps integration would be
> sufficient, but I think its a critical part of the project.(And building
> decent communities worked well for Dmoz, and DAMN [for a while])

Agreed.

> Memes seem like a very interesting, but difficult building block.  

Well, it's a convenience; I hate the term, actually. I just needed something
roughly equivalent to "idea" but more ideological. And, yes, drift is hard,
but I think it's actually easier to keep these straight than stuff like
"topic". By "meme" I mean very specific, very limited claims, almost always
expressible in a single sentence. Here are some:

 "Carlo Giuliani had a violent past"
 "Carlo Giuliani intended to throw a fire extinguisher into the jeep"
 "The carabiniera who shot Carlo Giuliani was terrified for his life"

and so on.

These arise *in* the primary sources, in the wire service reports, in the
reportage. Now, of course, editorial decisions will have to be made and
there will be drift and it's not a simple thing. But it's doable.

  Topics
> are hard enough to prevent "semantic drift" on, and those are usually
> simpler concepts.  Difficult but neat.  Seems like you'll want to have a
> hierarchy of these in order to simplify classification, and allow editors
> to browse for currently recognized memes.  

Absolutely. The point isn't so much to proliferate the number of them as
much as to track the instances of them and their mutation.

  As a bit of social engineering,
> I think one should emphasize all over the site how memes, are flexible,
> evolving concepts, and so people won't be surprised when a meme moves out
> from under them.

Yep.

> Not sure how Talking Points are directly related to the project.  Useful
> though.  I'm really sick of the "Inarticulate Protester" cliche.(Or is
> that a meme?)

It's a cliche and a meme. A cliche because it's part of the conventional
wisdom and is false.

The talking points idea is that once we've identified stable cliches, that
is, the atomic bits of the conventional wisdom, we want to -- here's the
exercising intellectual ability part -- provide responses to each one, in
the form of media kits and the like, so that protesters, their support
groups 'back home', friendly journos, and the like can knock them down and
we can do so in a kind of unified way.

I read just about every editorial published in English in the US about
Seattle last year. I'd say the number of atomic bits of the conventional
wisdom worth responding to was about 25 to 40 -- almost all of which I saw
in Quebec coverage, Genoa coverage, etc.

These things are very stable (at a high enough level; there are always
detail shifts and the like, but often of a sort that you can simply ignore).
And so it's possible to track, watch, and respond to them systematically.

Let's take one example: "Antiglobalization protests are a threat to
democracy". That's not only very consistent, and ubiquitous, it's also
*classic*, going back at least to Vietnam War and the social movements of
the 60s. It has variations of course: Bush's plaint that "protesters are
interfering with democratically-elected leaders" is a version of it.

It's utterly predictable and devastatingly answerable with a little
recitation on "the source of political authority rests with the sovereign,
who's consent to be governed can be revoked" etc.

> What is a pull quote?  Whatever it is, I would say mark it inline, but
> store a copy outside, cut down on processing.

It's a particularly striking or notable or quotable bit of a media report or
editorial. Marking them makes the job of responding more concrete since you
can quote what the wankers say back at them. (It also provides a way for an
editorialist like me to scan a page of newly archived articles, looking for
a quote that has the just-right nuance or language I want to use in some
piece I'm working on.)

> This duplication of data shouldn't be a problem because presumably the
> archived documents are static, and the data won't need to be kept in sync.

Plus, disk is cheap.

> Do we want to deal with the chase were the documents aren't static? (CNN
> has a really interesting and effective technique of writing stories
> gradually as a story unfolds, editing it every 30-45 minutes, might be
> interesting to catch that evolution)

That's hard, and I haven't thought about it much. AP/Reuters does something
similar, but in their case, they versions come out seriatim, so you can call
each one a "version" and track the evolution that way. I'm sure we'll try to
have support specifically for the wire services doing that. Since CNN
publishes new versions in situ, it's a lot harder to keep up. Though I guess
you could have a servlet poll a URI every n minutes, extract the <body>,
hash it, compare it to the last polled version, -- but that gets really
messy. Probably better to just assume -- and I think this may well be
warranted by the facts -- that CNN is tracking wire service changes.

> Do we need to get a normalized list of media organizations?  (Seems like
> one should exist somewhere) URIs could be sufficient, as long as we stick
> to web accessible resources.  

Absolutely, to both. (I think sticking to Web is perfectly reasonable too.)

  Definitely want to normalize the list so one
> can add neat info like a graph of media ownership to the analysis.(e.g.
> isn't it interesting that Fox News, and the SkyNews seem to always agree?)

Yes and yes. Good.

> What are the laws concerning archiving and making available this material?
> Especially if we're arching in digital format material which one is
> general required to pay for? (WSJ, on NYTimes archives)

Well, fair use applies to this project, I think. Plus, archiving is one
thing, publishing archival material is another, and, technically, we won't
be publishing it as much as allowing people to look at slices across the
archive.

I never read NYT because of the registration thing, but we'll have to sort
that out. It can't be ignored in this project.

> Webware seems cool, glad to see that python is starting to develop a
> coherent web development story.  I have to admit to being really puzzled
> by the appeal of AOLserver, Postgres, or OpenACS.  

The appeal of AOLserver is that it's way more robust and faster than Apache.
Also, the PyWX project means you can write AOLserver "servlets" in Python
(and you have always been able to write them in TCL; not my cup o'tea
though). PyWX exposes the entire AOLserver C API to Python. Pretty cool.

  Using Postgres has
> definitely cut into the number of people excited to work on active.(not as
> much as the quality of the active code base has, but its contributed)

Yeah, I just think it's more robust as an RDBMS than MySQL, but I also don't
have a strong opinion either way.

> Seems like AOLserver will have the same effect.  And OpenACS was still in
> TCL last time I checked.

Nah, since at the application level, people aren't writing AOLserver stuff
-- and may well not even have to know it's not Apache -- but, rather, we're
writing *Webware* apps, and so most if not all of AOLserver is just hidden.

As for OpenACS, yes, it's TCL, but there are TCL hackers in the world, and
you can call most of AOLserver (and hence OpenACS) TCL from PyWX Python. I
suspect we can find one or two TCL hackers to play along where needed (if at
all). Plus, the OpenACS thing is, well, open, so we might just steal some
good SQL data models and reimplement the forums and notification stuff in
Webware.

> Exciting

Glad you're interested.

Best,
Kendall
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.