Re: Installer Design Meeting, Tuesday the 5th (yeah tonight) at 9pm EST

Preston Cody <[email protected]> Wed, 7 Dec 2005 19:52:47 -0500
Newsgroups gmane.linux.gentoo.installer
Message-ID <[email protected]>
------=_Part_6866_21781586.1134003167647
Content-Type: text/plain; charset=ISO-8859-1
Content-Transfer-Encoding: quoted-printable
Content-Disposition: inline

A summary will be posted soon.  Be patient.
I'll attach an abbreviated logfile now for your amusement.
Enjoy,
-Codeman

On 12/7/05, Michiel de Bruijne <[email protected]> wrote:
> On Tuesday 06 December 2005 15:37, Preston Cody wrote:
> > Gentoo Installer Design Meeting, Tuesday the 5th (yeah tonight) at 9pm =
EST
> > in #gentoo-installer
>
> Hi, I saw this message too late, can someone post the IRC-log (or a point=
er)
> to this list?
>
> Thanks!
> --
> [email protected] mailing list
>
>

------=_Part_6866_21781586.1134003167647
Content-Type: application/octet-stream; name=meeting_log_12_06_05.log
Content-Transfer-Encoding: 7bit
Content-Disposition: attachment; filename="meeting_log_12_06_05.log"

19:52 <@codeman> ok sound off, who's here?
19:52 <@esammer> <---
19:52  * agaffney belches
19:52  * BenUrban is here
19:52 <@blackace> aye
19:52 <@AllanonJL> hello
19:53 <@blackace> we may need to wait another 7 or 8 minutes
19:53 <@codeman> eh i'll start talking status update stuff ahead of time since that's boring but should be done
19:54 <@codeman> so like, our last meeting was many months ago
19:54 <@codeman> and that was the 0.1 "alpha" days.
19:54 <@codeman> we had our nice alpha release
19:54 <@codeman> and it was full of bugs as expected
19:54 <@codeman> and then we fixed the bugs
19:54 <@codeman> and then told everyone to update to CVS
19:55 <@codeman> eventually we finally got 0.2 out with the -r1 release
19:55 <@codeman> and this appears to be much more stable (i.e. no fried partiton tables)
19:55 <@agaffney> hurray! :P
19:55 <@codeman> since then we've got a new FE basically ready to go, webgli
19:56 <@codeman> for those who have mercy on my desktop, http://24.149.145.141:8000/webgli/
19:56 <@codeman> if you want to see it in action
19:56  * agaffney watches gliserv.py explode
19:56  * BenUrban wonders how much ram is on that box
19:56 <@codeman> hey, don't repartition my desktop!
19:56 <+BenUrban> lmao
19:56 <@agaffney> delete! delete! delete!
19:57 <@codeman> anywhoo, recently i went through and made a nice comparison chart of the differences between the FEs
19:57 <@codeman> and then set about trying to make them all functionally identical to a certain extent
19:57 <@agaffney> http://dev.gentoo.org/~agaffney/gli/comparison.html should be up to date
19:58 <@codeman> it's my belief that the GTK FE is designed more for your average end-user and thus should take that into account
19:58 <@agaffney> well, it's not *designed* that way, but that's who it will end up being used by
19:58 <@codeman> so, for instance, skipping to and from steps is a *bad* thing in that FE
19:59 <@agaffney> it's really a bad thing in any FE
19:59 <@codeman> true.  it should maintain all of the functionality of all the other FEs
19:59 <@agaffney> since later steps can depend on the values from earlier ones
19:59 <@codeman> agaffney: well gli-dialog has a wizard mode, but at the end you can go back to any step to edit values
19:59 <@codeman> and yes, i know you can screw things up that way
20:00 <@agaffney> yeah, that's a Bad Thing(TM)
20:00 <@agaffney> what you said :P
20:00 <@codeman> but very usefull because it's hard to undo in dialog
20:00 <@blackace> let's move this to "problems with the current FEs" guys
20:00 <@codeman> right.
20:01 <@codeman> so webgli was designed to be kindof a plugin to a higher-up application that would deal with the second of our main goals of GLI
20:01 <@codeman> being mass-deployment
20:02 <@codeman> ok, so agaffney and i threw together something i called mastergli, which is badly in need of a new name
20:02 -!- npmccallum-work [n=npmccall-xC5Qu3Aly4ifc8rMBmLxT2ERkLen7jch6yr7iItg/[email protected]] has joined #gentoo-installer
20:02 <@codeman> http://24.149.145.141:8000/
20:02 <@codeman> basically gliserv runs both
20:02 <+npmccallum-work> hola
20:02 <@agaffney> npmccallum-work: the party's already in full swing
20:02 <@codeman> webgli is a subdirectory and hence acts like a module of mastergli
20:03  * blackace is gonna kickban gibot and +q/-v anyone who interrupts codeman again.
20:03 <+antarus|osx> curse words on the front page make a great presentation :P
20:03  * codeman points at agaffney
20:03 <@agaffney> codeman: pfft, you can easily change that
20:03 <@agaffney> I put that in there when I was writing it locally before I committed it
20:03 <@codeman> ok, so basically the main point of "mastergli"... ok wait.. lets stop here and come up with a better name right now
20:04 <@blackace> gliid
20:04 <@codeman> because if i keep using it it'll stick
20:04 <@blackace> gli install daemon
20:04 <+antarus|osx> glid
20:04 <+antarus|osx> the i in gli is for install :p
20:04 <+BenUrban> gli-server?
20:04 <@blackace> installer
20:04 <@agaffney> gliiiiiiiiiiiiiiiiiid
20:04 <@codeman> antarus|osx: glid is what i use to refer to gli-dialog
20:04 <@blackace> see?  gliid is catchy
20:04 <@codeman> if we could add an e to the end it could be glide
20:04 <+BenUrban> iirc that already exists
20:05 <+BenUrban> err maybe not
20:05 <+antarus|osx> Gentoo Linux Installer Daemon
20:05 <@codeman> so far i like gliid the best
20:05  * BenUrban thinks it's too easy to typo gliid/glid
20:06 <@agaffney> Gentoo Linux Installer Install Daemon
20:06 <@agaffney> meh
20:06 <@AllanonJL> i think this is a moot point that can be discussed later :)
20:06 <+antarus|osx> can we perhaps vote on this later? :p
20:06 <@codeman> ok
20:06 <+wolf31o2> Gentoo Linux Installer Install Daemon... brought to you by the Department of Redundqancy Department
20:06 <@codeman> i'll call it the Installer Daemon for now
20:06 <@codeman> since that describes it well
20:06 <@blackace> another option is glimd, Gentoo Linux Installer Management Daemon
20:06 <+wolf31o2> ooh... go blackace
20:07 <@blackace> it's purpose is not to install, but to manage installs and profiles.
20:07 <@codeman> so the basic idea is that clients netboot up (or boot a livecd)
20:07 <@agaffney> blackace: good point...I like it
20:07 <@agaffney> heh
20:07  * agaffney shuts up so codeman can continue
20:07 <@codeman> and then on bootup the client will reach out and touch the server and provide the Installer Daemon the networking info of that machine
20:08 <+antarus|osx> codeman, is there anything specifically tying them to netbooting/liveCDs?  I could run the software on say, a debian install to install on another hd fex?
20:08 <@codeman> specifically the MAC address
20:08 <@codeman> antarus|osx: not at all, just any environement with the appropriate dependencies and the client
20:08 -!- samyron [n=samyron-TDmef5qNbhSD0RbGiatluDonLGFlP/VaNCKqTw2Vcvvvt0rt8C/[email protected]] has joined #gentoo-installer
20:08  * antarus|osx thought so, please continue ;0
20:08 <@agaffney> antarus|osx: only the installer's dependencies...for the network client, probably just parted/pyparted
20:08 <@blackace> samyron: welcome, codeman is doing a status report.
20:09 <@codeman> agaffney: am i missing anything on the info provided?
20:09 <@samyron> blackace: tnx
20:09 <@codeman> why don't you talk about how it works
20:09 <+wolf31o2> by the way... I've got to run upstairs for a bit, but this is something I found pretty on-topic from GWN... (sorry for the interruption but I'm not sure when I'll be back)
20:09 <+wolf31o2> "One thing I think Gentoo needs is its own version of the SystemImager tool. A 'GentooImager' ebuild, if you will. SystemImager is tailored to Redhat/Fedora, and includes the whole kitchen sink in its distribution. A GentooImager ebuild, perhaps based on and borrowing from SystemImager, could just make the necessary tools -- dhcp, syslinux, tftp, etc. -- be dependencies. Beyond that, there's the matter of providing an environment
20:09 <+antarus|osx> wolf31o2, the 3x3 guy?
20:09 <+wolf31o2> for pxeboot to boot into and bootstrap a system and the other details of making system images. Anyway, its something to think about. When/if we build another hyperwall, I may start working on such a package. We'll see. I'm interested in hearing from others who have installed and maintained Gentoo on a cluster."
20:09 <@agaffney> client CD/net boots, client starts, broadcasts for server on UDP port 8001, server responds, client connects to server via XMLRPC now that it knows the IP, registers itself, and waits
20:09 <+wolf31o2> yes... perhaps he would be a good guy to contact
20:10 <@agaffney> the client polls the server every few seconds asking if it can start the install
20:10 <@codeman> wolf31o2: i'll get to that part later
20:11 <@agaffney> when the server says yes (triggered by the admin setting a flag in the web interface), the client downloads its client_profile and install_profile, and starts its install
20:11 <@agaffney> for each step completed, it calls a function via XMLRPC that updates its install status
20:11 <@agaffney> so the server knows where each client is
20:11 <+wolf31o2> I still hate the idea of broadcasting to find the server... it isn't very useful in places with small broadcast domains and multiple subnets
20:12 <@agaffney> wolf31o2: you can specify the server IP on the commandline when the client starts
20:12 <@agaffney> wolf31o2: if you specify it, it doesn't broadcast
20:12 <+wolf31o2> still doesn't change that a broadcast sucks
20:12 <@agaffney> what's the alternative?
20:12 <+antarus|osx> the netboot root fs should be able to have a IP hardcoded don't you think?
20:12 <@agaffney> antarus|osx: depends on if it's a "stock" netboot image or something custom
20:13 <+wolf31o2> well... the image to boot from has to come from somewhere... so the server already knows where it is
20:13 <+antarus|osx> but that only makes sense if your netbooting server is the install server or not :)
20:13 <@agaffney> of course, if the clients are netbooting, it could be passed as a kernel parameter
20:13 -!- wolf31o2|mobile [n=wolf31o2@gentoo/developer/wolf31o2] has joined #gentoo-installer
20:13 -!- mode/#gentoo-installer [+v wolf31o2|mobile] by ChanServ
20:13 <@agaffney> wolf31o2: also, if you CD boot the clients, they don't know where the server is :)
20:14 <+antarus|osx> I'd assume most large shops would make their owno CD
20:14 <+wolf31o2|mobile> I still think it is best to just rely on kernel boot options
20:14 <@agaffney> wolf31o2|mobile: what does broadcasting hurt?
20:14 <@agaffney> if it doesn't find it, it doesn't find it
20:14 <+antarus|osx> can you netboot a machine without a tftp server in the local subnet?
20:14 <@agaffney> yes
20:14 <+wolf31o2|mobile> antarus|osx: yes
20:14 <@codeman> guys, this discussion should be taken post-meeting
20:15 <@agaffney> agreed, let's continue
20:15  * antarus|osx adds netbooting to post meeting notes
20:15 <+antarus|osx> continue ;0
20:15 <@codeman> ok
20:15 <@codeman> so after talking with eric and another coworker who does a lot of the imaging where I work I see a lot of places that GLI can go
20:15 <@agaffney> anyone have any questions about the basics of the client/server relationship and the network install process?
20:16 <+wolf31o2|mobile> it's all pull, right... client pulls form server... no pushing
20:17 <+wolf31o2|mobile> s/form/from/
20:17 <@codeman> yes
20:17 -!- wolf31o2 [n=wolf31o2@gentoo/developer/wolf31o2] has quit [Read error: 113 (No route to host)]
20:17 -!- wolf31o2|mobile is now known as wolf31o2
20:18 <@agaffney> wolf31o2: right, otherwise the client has to run a XMLRPC server as well
20:18 <@codeman> ok, so one of the major problems with using the Installer Daemon for mass installs is it simply can't work for lots of machines due to a couple factors
20:18 <@codeman> most obvious is if you had every machine download and sync a portage tree, you'd get banned from a lot of sync servers
20:19  * BenUrban wonders if something could be rigged in the bitttorrent style
20:19 <@codeman> then there's the issue of customization into groups of profiles
20:19 <@agaffney> codeman: that's why you'd be smart and use a local rsync server or a local snapshot
20:19 <@codeman> BenUrban: well the simple solution to that is a common NFS mounted tree... but i'm sure antarus will have somethign to say about remote trees soon
20:19 <@AllanonJL> codeman: i think another big issue that it really should be using an image, and not every box "installing" from scratch
20:20 <@AllanonJL> then you can just multicast the image out
20:20 <@codeman> AllanonJL: EXACTLY :)
20:20 <+BenUrban> ideally the install daemon would give the option of automatically setting that up
20:20 <@agaffney> AllanonJL: we'll have that *also*
20:20 <+BenUrban> yeah
20:20 <@agaffney> AllanonJL: pretty much a stage4 install that just does partitioning and bootloader
20:20 <@blackace> which brings up the problem of machine type variances and how to handle that in a sane way
20:20 <@codeman> most sysadmins install images, not distributions.. be it windows or linux
20:21 <@codeman> so this brings us to the concept of "roles" or "modes" for the installer to run in
20:21 <@agaffney> blackace: I'd say that's up to the admin to handle...we'll provide support for specifying relative partition sizes
20:21 <@agaffney> not much else we can do
20:21  * antarus|osx thinks a lot of that is based on what you end up doing per 'node'
20:21 <@blackace> agaffney: build on the first machine of a type, then stage4 the others of that same type...you have to handle that somehow
20:22 <@codeman> we'll need to be able to install to a chroot so that the user can customize their image and then we'll need to be able to install an image (i.e. stage4) where only partitioning and bootloader takes place
20:22 <+antarus|osx> Partially I think it's due to the fact that this isn't an Installer anymore ;)
20:22 <+wolf31o2> ok... so you're talking a stage4, not an "image"
20:22 <@codeman> so agaffney suggested another layer on top of the Architecture templates called modes
20:22 <+BenUrban> wolf31o2: what's the diff?
20:22 <@codeman> where you could define the steps
20:22 <@blackace> BenUrban: stage4 == filesystem contents, image == disk image
20:23 <+BenUrban> oh
20:23 <@agaffney> wolf31o2: partition, extract stage4, install the bootloader, done
20:23 <+antarus|osx> I think by 'image' it's either a tarball or some dd deal, both make sense depending on your circumstances
20:23 <+wolf31o2> BenUrban: images are pretty damn useless on Linux unless you build an imaging tool that can understand the underlying filesystems for any that the admin might want to use
20:23 <+wolf31o2> agaffney: yeah... that's a good plan
20:23 <@codeman> agaffney: you skipped configure networking
20:23 <+BenUrban> well i assumed you were talking about filesystem images
20:23 <@blackace> floor back to codeman please.
20:23 <@codeman> ok so back to modes
20:23 <@codeman> how do we want to do this?
20:24  * antarus|osx likes roles better personally
20:24 <@codeman> we could define it right in the ArchitectureTemplate
20:24 <@AllanonJL> so is this like machine "Profiles" ?
20:24 <@codeman> and be forced to a step list that is architecture independent
20:24 <@agaffney> AllanonJL: no, full install vs. "stage4"
20:24 <@agaffney> AllanonJL: vs. "intall into chroot"
20:25 -!- bedros [n=bedros-ZcXrCSuhvln8CQt26AM2xgur6w5192QaFQehaosVqYBHxeISYlDBzl6hYfS7NtTn@public.gmane.org] has joined #gentoo-installer
20:25 <@codeman> or we could add another layer of templates that define the roles and do something like ChrootTemplate imports amd64 template imports ArchTemplate
20:25 <@codeman> esammer: anything to say?
20:26 <@blackace> codeman: I envision being able to build an installprofile with "selectors" (think css) that can select things from the ArchTemplate.
20:26 <@agaffney> blackace: eh? I don't get what you're saying
20:26 <@codeman> i don't want people being able to pick and choose steps
20:26 <+BenUrban> blackace: you mean like selectively excluding steps?
20:26 <@codeman> blackace: you can already skip most steps by leaving blank values
20:27 <@blackace> like: useflags[arch="nocona"] { different use flags; }
20:27 <@esammer> i feel like i'd have to think about it for some time before being able to say anything definitive. i think there's a few things to consider, but more so than anything, it's important to understand the goal before talking about data structures and how to implement something like this.
20:28  * BenUrban thinks that that sounded awfully generic
20:28 <@codeman> ok so the goal is to be able to selectively use the backend
20:28 <@agaffney> esammer: unfortunately, the 2 lead devs (codeman and I) have a tendency to jump right into things :)
20:28 <@esammer> BenUrban: i know. it's a generic approach to design.
20:28 <@codeman> because it can do a full install, we know that.  but for what we're talking about it'll need to become almost modular
20:28 <@blackace> codeman: it's either that, or define groups of machines server-side and assign them all different installprofiles
20:28 <+antarus|osx> Ok so you want to use the installer backend to "install" the 'system image'/'stage4'?
20:29 <@codeman> antarus|osx: yes
20:29 <@esammer> agaffney: understandably so. it's nice to produce code that works, but it's nicer to produce code that will work tomorrow as well.
20:29 <@agaffney> esammer: indeed
20:29 <@esammer> agaffney: you need to start somewhere, but it's good to know where to start.
20:30 <@esammer> more than anything, it's imperitive that everyone be on the same page.
20:30 <@agaffney> esammer: yes dad
20:30 <@blackace> :)
20:30 <@esammer> right on. so, that's it for me.
20:30 <@codeman> ok moving on
20:30 <+antarus|osx> codeman, so assuming some local client that will do the install in a chroot, obviously you need to skip steps, and thats not easy in the current implementation?
20:31 <+antarus|osx> or it' s just not modular enough?
20:31 <@agaffney> esammer: going to bed already?
20:31 <@esammer> agaffney: nope. getting fed up already.
20:31 <@codeman> antarus|osx: no it's easy, but it just needs a different list of steps
20:32 <@agaffney> antarus|osx: I've already been thinking about ways to do that for local chroot installs
20:32  * blackace notes everyone needs to STFU so codeman can finish the status report that is now 40 minutes in.
20:32 <@codeman> MOVING ON, so after a nice long conversation with the image installer, he uses the program wolf mentioned called SystemInstaller
20:32 <@codeman> it's got a pretty crappy sourceforge page
20:32 <@codeman> so there's no eye candy
20:32 <@codeman> but it's similar in function to m23 (which has a nice site)
20:33 <@codeman> basically it deals with installing compressed images to machines
20:34 <@codeman> it does a lot of stuff i don't personally think is very smart 
20:34 <@agaffney> codeman: dd'd images?
20:34 <@codeman> esammer: do you know?
20:34 <@codeman> i don't know which method it uses to actually install the image
20:34 <@esammer> it's a bit more than that.
20:35 <+antarus|osx> implementation aside, what does it accomplish?
20:35 <@esammer> you need to boot, "install", post configure (networking, boot loader, services based on a profile, init scripts, packages in our case debian, but can be anything, etc.)
20:36 -!- npmccallum-work [n=npmccall-xC5Qu3Aly4ifc8rMBmLxT2ERkLen7jch6yr7iItg/[email protected]] has quit ["Leaving"]
20:36 <@codeman> esammer: focusing on the "install" tho
20:36 <@esammer> you need to generate host keys, set PKI info, create distribution queues, update the rest of the machines in teh cluster, etc.
20:36 <@codeman> say what?
20:37 <@agaffney> for any post-install setup functionality that our installer is lacking, the installer does support a post-install script
20:37 <@esammer> the install aspect is primarily an untar, yes. ala installation with a stage3
20:37 <@codeman> esammer: you lost me at distribution queues
20:37 <+antarus|osx> most of that is handled during the install itself, PKI keys are probably custom post-install script, along with what I'd call site dependent things
20:37 <+antarus|osx> unless you can standardize each task ( ala distribution queues ;) )
20:38 <@esammer> well, if i may explain...
20:38 <@agaffney> codeman: cluster stuff, I assume
20:38 <@agaffney> esammer: please do
20:38 <+antarus|osx> esammer, if you can devulge also, your usage of machines..cluster, desktops..etc?
20:38 <@esammer> so when i roll out a box for deployment, it usually needs to receive remote code updates or receive ssh connections from automated processes, have log files picked up for backend processing, etc.
20:40 <@esammer> antarus|osx: load balanced sets of "cloned" machines (loose cluster), processing distributed (non-connected) nodes for reporting, banks of independent machines (i.e. outgoing MTAs), etc.
20:40 <@agaffney> keep in mind that the installer should really only go as far as the handbook does...any further post-install configuration should be done with the use of a post-install script
20:40 <@agaffney> anything more than that and the installer starts to get *very* complicated
20:40 <@blackace> so, emerge some stuff, grab some ssh keys, and modify a syslog config...sounds to me like a job for a post-install script maybe with some cfengine-foo.
20:40 <@agaffney> blackace: exactly
20:40 <@blackace> not GLI's job.
20:40 <@codeman> agaffney: this whole damn meeting is intended to talk about things BEYOND GLI
20:41 <@agaffney> codeman: if you mean things that GLI can do in the future, I still think the current topic is beyond GLI
20:41 <@esammer> like i said, it's about goals. i'm just explaining how something is used in a real environment full of machines that are rolled out fast.
20:41 <@codeman> i.e. how to install to chroot, allow for customization, and then send off an image to a series of machines
20:41 <+antarus|osx> esammer, the point being this is all done by SystemInstaller?
20:41 <@blackace> codeman: ok, but not site-specific customizations, correct?
20:41 <@codeman> and then use a central utility to update them, get logs, etc
20:42 -!- npmccallum-work [n=npmccall-xC5Qu3Aly4ifc8rMBmLxT2ERkLen7jch6yr7iItg/[email protected]] has joined #gentoo-installer
20:42 <@esammer> antarus|osx: parts of it, but it is limited, which is the point.
20:42  * antarus|osx nods
20:42 <+antarus|osx> ok then
20:42 <@esammer> antarus|osx: more than SystemImager is needed.
20:42 <+antarus|osx> esammer, so your point being, no one wants a thing that just installs machines, you want the whole shebang?
20:42 <@esammer> installation shouldn't be considered a one time thing. it bleeds into maintenance and configuration management.
20:42 <@codeman> from what i talked to john case, he would like to be able to define a profile for a group of machines
20:43 <@codeman> if i may.  one sec.
20:43 <@AllanonJL> codeman: while you are at it, want to tackle policy enforcement? because that is the next step
20:43 <@blackace> let's at least agree to leave logging out of it, central logging is something builtin to syslog-ng
20:43 <+antarus|osx> esammer, I thought as much.
20:44 <@codeman> once the profile can be defined (a normal GLI one) and then use that profile to install to a chroot, the user can then customize that install to their hearts content.  This becomes the "lead" machine of the group.
20:44 <@codeman> once that is done, then using a "stage4" role, the same profile can be used in conjunction with the "lead" machine to install to any number of machines
20:45 <@blackace> any reason we can't split this out to a separate "replication" backend?
20:45 <@codeman> so basically one profile/group per type of machine, be it MTAs or Printerservers
20:45 <@codeman> blackace: that's what i'm talking about
20:45 <@agaffney> codeman: so the "lead" would be (or atleast should be) a staging box (during and post-install)?
20:45 <@codeman> yes.
20:45 <+wolf31o2> just curious... but why not utilize catalyst? you can use a stage4 as a seed for a stage4... the spec file would be the "profile" along with an overlay... might save a ton of time
20:46 <@agaffney> so post-install, this machine would get an update and then send it out to the other members of the group via a binary?
20:46 <@codeman> it would also be the machine you would try updates on first before sending them out to the rest of the group
20:46 <@blackace> codeman: ok, I thought you meant this to be part of what we already have.
20:46 <+antarus|osx> s/binary/something/
20:46 <@agaffney> codeman: that's what I meant by "staging box"
20:46 <@codeman> agaffney: i know :)
20:47 <@codeman> ok another feature john case said was a must-have was the ability to rollback an update.  this, however, may be more of a portage issue than anything i could do
20:47 <@codeman> s/could/should
20:47  * antarus|osx coughs softly
20:47 <@codeman> antarus|osx: speak
20:47 <+antarus|osx> Depends on your timeframe
20:47 <@codeman> a brief overview of what's coming up in portage.
20:47 <@agaffney> antarus|osx: no specific timeframe
20:47 <@codeman> please
20:48 <+antarus|osx> ok
20:48 <+antarus|osx> FYI, there is no prediction for the Savioro branch to be done
20:48 <+antarus|osx> wow, bad mispelling ;)  So Savior is all about user power, you want thing X, you just need to derive from Framework classes and implement it
20:49 <@codeman> can you dumb that down a bit
20:49 <@agaffney> codeman: everything will be very OO and expandable
20:49 <@codeman> thank you
20:49 <+antarus|osx> heh ;)
20:50 <@codeman> what are some features planned?
20:50 <@agaffney> instead of the mish mash of procedural evilness that portage is now
20:50 <+antarus|osx> Remote Tree, Remote VDB, SQL cache backends,
20:50 <+antarus|osx> Portage Daemon
20:51 <@agaffney> remote tree == something other than flatfile?
20:51 <+antarus|osx> agaffney, whatever you  like, as long as you write the code to do it
20:51 <@agaffney> k
20:51 <+antarus|osx> agaffney, and it adheres to the class interface
20:51 <+antarus|osx> So for example, rollbacks
20:52 <@codeman> antarus|osx: describe portage daemon
20:52 <+antarus|osx> Tricky buggers, depends on how reliable, obviously if you hose glibc or something no rollback software will save you ;)
20:52 <+antarus|osx> Basically a portaged that listens to both remote and local commands.
20:52 <@agaffney> antarus|osx: what about a static 'savemya$$' program?
20:53 <@agaffney> codeman: basically, you could get a deptree for 'world' on box B from box A
20:53 <@blackace> following the "lead box" model, the management server could stage4 it before rolling changes out to it and provide the stage4 as a rollback option.
20:54 <+antarus|osx> blackace, ehhh messy, especially if nodes are different
20:54 <@agaffney> blackace: seems a bit drastic
20:54 <+antarus|osx> I was thinking portaged would backup locally.
20:54 <@blackace> antarus|osx: talking just the "lead" box here.
20:54 <+antarus|osx> since we already know what files we install, just back them up beforehand
20:54 <@codeman> blackace: that woudln't work.  you'd clobber things like the custom networking settings on the other machines
20:54 <@agaffney> antarus|osx: what about files that are generated in pkg_postinst() or by the user?
20:55 <+antarus|osx> but then it's how long to do you keep backups for...how long are rollbacks optional?
20:55 <@blackace> codeman: just the "lead" box since you roll out updates to it first to test them before rolling them out to the other machines.
20:55 <+antarus|osx> agaffney, "by the user" meaning config files?
20:55 <@agaffney> antarus|osx: primarily
20:55 <+antarus|osx> I mean if you upgraded mysql and hosed your db I don't see how it's portage's job to get you back to sq 1 ;)
20:55 <@blackace> codeman: this can be in addition to whatever the portage team comes up with in Savior, and could then still be useful in case cp, tar, etc. stops working.
20:56 <@codeman> i was hoping the future portage could be able to keep track of what's been updated and be able to just remerge the old version back
20:56 <@codeman> blackace: very true
20:56 <@agaffney> codeman: it would be relatively easy to write a quick util that does that from the emerge.log
20:56 <@agaffney> codeman: right now
20:56 <@codeman> agaffney: ok lets do it that way for now then
20:57 <@codeman> it seems like if we lay out a nice framework for a higher up util, by the time we get it functional portage 3.x will be ready for us to utilize
20:57  * antarus|osx notes that implementation of rollback is a whole discussion in itself
20:58 <@blackace> if I suspect something may break, I typically quickpkg first, then upgrade, then I can unmerge the broken package, and remerge the old one.
20:58 <@blackace> this can all be handled by the management server.
20:58 <@agaffney> blackace: as do I...I assume the rollback would automate this process (and the restoration)
20:58 <@blackace> agaffney: right
20:58 <+antarus|osx> I think the big thing here is that portage doesn't care what the underlying information looks like, how it's stored, where it's stored, nothing, as long as you give it what the interface needs.
20:59 <@agaffney> right, with portage, it's "automatic" and flexible
20:59 <+antarus|osx> I am not sure when this will be done, I think the bigger hope is to finish the generic classes and framework, and then hope that more people contribute derived stuff
20:59 <@esammer> antarus|osx: that's pretty exciting. i haven't been paying attention to portage development, but it sounds like you've done quite some work.
20:59 <+antarus|osx> s/you've/ferringb/
21:00 <@agaffney> esammer: portage svn looks just a bit different from portage stable :)
21:00 <@esammer> fair enough.
21:00 <+antarus|osx> I haven't done squat, I'm merely their evangelist ;)
21:00 <@codeman> this higherup project will most definitatly have to work a lot more closely with portage, just like we've been snuggling up to wolf31o2 to get GLI out.
21:00 <@codeman> so let me share with you guys a mockup i made really quickly a few days ago after having these discussions w/ people.  http://24.149.145.141/GLI/glisysau.htm
21:01  * codeman often has fun w/ dreamweaver on the train ride home
21:01 <@codeman> so how can ^^ be changed/improved?
21:01 <@codeman> what cool new things should it do?
21:02 <+wolf31o2> anybody seen RHN Satellite?
21:02 <+antarus|osx> s/lead/staging
21:02 <@agaffney> wolf31o2: never heard of it
21:02 <@blackace> first of all, GLI != this new thing, say GLR, Gentoo Linux Replicator...seems to me it should be separate from GLI and maybe doesn't belong in releng, but maybe server
21:03 <@agaffney> blackace: I'm leaning towards that as well
21:03 <@agaffney> it seems different enough from the installer that it's a stretch to try to integrate them
21:03 <+wolf31o2> agaffney: it is a local copy of RHN for RHEL... it is used for provisioning, maintenance and now even monitoring of servers
21:03 <@agaffney> wolf31o2: I've never used RHEL and I haven't touched RH since 8.1
21:04 <@blackace> doesn't stop them from working hand in hand, but they should definitely not be integrated such that they cannot be used apart from each other
21:04 <@codeman> blackace: this is definitatly not GLI.. this is a higherup application that uses GLI like a module, to get installs done.
21:04 <+wolf31o2> agaffney: the point is that it is a consistent interface for doing these things... something we really should be doing... only better
21:04 <@blackace> codeman: cool
21:04 <@blackace> codeman: or maybe they're both modules for this GLIMD?
21:04 <@codeman> and it's exactly where gentoo needs to go if it wants to be nice in the E word (enterprise).
21:04 <@agaffney> blackace: I think that's what he was suggesting
21:05 <@blackace> or maybe a better name would be Gentoo Linux Deployment Management Daemon
21:05 <@esammer> i need to get away from a computer for the remaining hour before i sleep. it's nice to see everyone again.
21:05 <+wolf31o2> or just GLMD... since it isn't just deployment... it's everything
21:05 <+wolf31o2> esammer: same, man
21:05 <@agaffney> esammer: good night and thanks for your input
21:05 <@codeman> later eric
21:05 <@blackace> esammer: thanks for being here :)
21:05 <@blackace> wolf31o2: yeah, true
21:05 <@AllanonJL> codeman: perhaps a name for the management (configuration) prog could be Gentoo Operations Manager ?
21:05 <@esammer> i'm sure i will see you all again.
21:05 <@esammer> soon.
21:06 <+wolf31o2> anyway... name isn't important... design/functionality is
21:06 <+wolf31o2> heh
21:06 <@agaffney> esammer: we'll post a log of all this if you want to read it later
21:06 <@esammer> agaffney: sounds good.
21:07 <@AllanonJL> i think we need to define the scope of the different projects, basically defining where one project starts and another ends, and then defining where the interfaces have to be...ideas, comments?
21:07 <@codeman> ok so ignoring the name issue again for now
21:08 <@codeman> well the highest level i think should just be linking the lower-level projects
21:08 <@blackace> AllanonJL: GLMD in the server project, uses optional components GLR in the server project, and GLI in the releng project.
21:08 <+wolf31o2> well... I would think the idea would be to determine what kind of a daemon do we really need/want... what can it do? what should it do?
21:09 <+wolf31o2> I mean... is it supposed to be a "hand of god" type of daemon where you can control *everything* from a single console somewhere?
21:09 <@agaffney> wolf31o2: gliserv (the web/xmlrpc/udp ping server) can be used as a general purpose daemon for this
21:09 <@codeman> wolf31o2: i invision it getting there in small steps
21:09 <+wolf31o2> or are we limiting the scope
21:09 <@agaffney> wolf31o2: nobody is quite sure about that yet :)
21:09 <+wolf31o2> for example, with what you have on your mockup, it is basically provisioning... with the "System Updater" being maintenance
21:10 <@codeman> hrm.. you have a point
21:10 <@codeman> that doesn't really work out too well.. because the maintenance is the majority of the work it'd ahve to do
21:11 <@blackace> I would suggest here the possibility of something like xmpp instead of xmlrpc...something that would allow two-way communication so GLMD can talk to GLI and GLR "plugins" (which in the case of GLI is basically a headless frontend, just even more so than the current webgli concept)...GLMD would be the head of all the different beasts.
21:11 <@codeman> so the System Updater really needs to be many separate things
21:11 <+wolf31o2> would you guys like for me to explain what satellite does? it might help you understand what at least "the competition" is offering
21:11 <@codeman> wolf31o2: yes please
21:11 <@agaffney> blackace: what's xmpp? 2-way can be done with xmlrpc if the "client" runs a xmlrpc daemon as well
21:11 <@codeman> URL's with pics would rock if you ahve em
21:11 <+wolf31o2> blackace: you can't guarantee two-way communication like that... I'd say no
21:11 -!- samyron [n=samyron-TDmef5qNbhSD0RbGiatluDonLGFlP/VaNCKqTw2Vcvvvt0rt8C/[email protected]] has quit []
21:11 <+wolf31o2> codeman: I don't
21:11 <@blackace> wolf31o2: how do you mean?
21:12 <@blackace> agaffney: the stuff jabber uses to communicate
21:12 <@blackace> agaffney: it's xml based, but multi-way and addressable
21:12 <@agaffney> blackace: how's that different than both the client and server running xmlrpc daemons?
21:12 <@agaffney> blackace: seems like more than we need
21:12 <+wolf31o2> anyway... satellite is broken up into a few parts
21:13 <@blackace> agaffney: one server, many clients...it doesn't matter, simple IPC would work, just the interface needs to be able to talk to headless FEs and they back.
21:13 <@blackace> wolf31o2: go ahead :)
21:13 <+wolf31o2> provisioning - this is kickstart... it has the concept of a kickstart profile, which defines a server type... this tells it what channels to subscribe the server to... it also includes any filesystem overlays (yeah, rh does it this way) and pre and post install scripts...
21:14 <+wolf31o2> a "channel" in rhn is just a software collection... a server can belong to any number of channels
21:14 <@blackace> like a portage overlay in other words?
21:14 <+wolf31o2> blackace: something like that, yeah
21:15 <+wolf31o2> ok... once the machine is provisioned, it also has a daemon that checks in every so often with the server... this is how the server knows when machines are out of date... it also acts as a job scheduler... so you can say "update all packages on this server" and next time it checks in, it does
21:15 <+wolf31o2> the client actually does the calculation, based on what is available on the server
21:16 <+wolf31o2> so the easiest way to think of it is the server houses the tree, plus any overlays (channels)... the client would do an "emerge -vuDN world" and give the output to the server
21:16 <+wolf31o2> so now the server knows what is out of date on the client... you can tell it to update everything or a single package...
21:17 <+wolf31o2> now... the thing with satellite is it is *not* an actual control tool
21:17 <+wolf31o2> once a box is provisioned, you have to get on the box to change anything... all it does it package management
21:17 <+wolf31o2> it has no rollback, either
21:18 <@codeman> well we can pwn that then!
21:18 <+wolf31o2> so we'd already be at an advantage here... the one thing is that it does keep all of the packages (until you clean them out) even the older ones
21:18 <@codeman> wolf31o2: so it does sound a bit to me like m23 again, but for RH instead of debian
21:18 <@codeman> http://m23.sf.net  has nice eye candy
21:19 <+wolf31o2> now... using portage, this would be simple... keep the packages on the "server" box... need to rollback, BINHOST="http://server/channel" emerge --oneshot -K =cat/pkg-version
21:20 <@codeman> and as far as defining packages each group could just have its own overlay
21:20 <+wolf31o2> right... making it easy as hell to export as a binhost/nfs export
21:20 <@blackace> right, sync multiple overlays
21:20 <+wolf31o2> remote overlays would really own this... but yeah
21:20 <+wolf31o2> err... remote portage trees, I mean
21:21 <+wolf31o2> but for the short term, syncing multiple overlays would work fine
21:22 <+wolf31o2> the idea is you have a client-side daemon/cron type thing that runs every so often... then a way to also do it from the cli *now*
21:22 <+wolf31o2> basically... have multiple "entitlements" (RH does this too) or whatever you want to call it
21:22 <+wolf31o2> like... "provisioning"
21:22 <@codeman> fancy words for "install" eh?
21:22 <+wolf31o2> means it builds via GLMD/GLI, etc... then after that, it is a "normal" Gentoo box
21:23 <@blackace> this can be done with cfengine now, but cfengine isn't exactly user friendly
21:23 <+wolf31o2> emerge sync, portage, no special channels, etc
21:23 <+wolf31o2> "management" would be all of provisioning, but tied to the server more closely... syncs are tied to what is on the server/channels/etc
21:24 <+wolf31o2> so "management" would be more for enterprise use
21:24 <+wolf31o2> "provisioning" would be for deployment
21:25 <+wolf31o2> or just mass installs/setup
21:25 <+wolf31o2> so somebody would use provisioning if they were... say... Genesi... for setting up tons of boxes, then shipping them to customers
21:25 <+wolf31o2> whereas somebody making a cluster would totally do management
21:26 <@codeman> so install and then just leave it with a little admin tool in the cron that does all the work?
21:26 <+wolf31o2> or an enterprise
21:26 <+wolf31o2> codeman: basically, yeah...
21:26 <@codeman> makes sense
21:26 <@blackace> codeman: a replication client, and/or a management client
21:27 <@blackace> seems to me it makes sense at this point to have three separate clients like that, install, replicate, and manage
21:27 <@codeman> well you get into the issue there of which machine's in charge, the client or the server?  does the server "push out" updates
21:27 <+wolf31o2> yeah... except in the case of replication (like my genesi example) they can just remove the client/cron job before shipping it out and it is a normal Gentoo box
21:27 <@blackace> right
21:27 <+wolf31o2> codeman: no... the client pulls them
21:28 <@blackace> whereas people running some diverse clusters would leave the replication client so they could pick an arbitrary box and replicate it when they add a node
21:28 <@codeman> but how does the client know what to pull?
21:28 <@blackace> it asks the server
21:28 <@agaffney> codeman: it asks the server
21:28 <+wolf31o2> codeman: basically, the server is just a repository... you can use the server to schedule stuff for the client... but the client is really the "boss"
21:28 <+wolf31o2> right
21:28 <+wolf31o2> like... ok... you could do: update world
21:29 <+wolf31o2> and it would update everything with the latest on the server
21:29 <+wolf31o2> or you could do; update catalyst
21:29 <+wolf31o2> and it would only update catalyst
21:29 <+wolf31o2> the server would still show any other packages as needing updates
21:29 <@codeman> yah but the client's gonna need some special information to distingush itself
21:29 <+wolf31o2> so?
21:29 <@codeman> so that my very-stable boxes don't get updates or whatever
21:29 <@agaffney> codeman: MAC address?
21:30 <@agaffney> we already figured this one out :)
21:30 <@codeman> i was thinking htat.  and stored on the server
21:30 <@codeman> so the server's definitatly gonna need a database of client info
21:30 <+wolf31o2> MAC address could be used... GUID could be used... personally, I would go with a completely unique ID created during provisioning and stored on the client
21:30 <@agaffney> already got that covered with GLIServerProfile
21:30 <@agaffney> although, it would need to be extended
21:30 <@blackace> codeman: of course, otherwise how does the server know what to tell to a client when the client asks "yo, whussup?"
21:30 <@codeman> a good bit
21:31 <+wolf31o2> because... for example... if a box/NIC died, you could just copy it to new hardware (or replace the NIC) and not lose the client's ID
21:31 <@codeman> blackace: well the client could say "yo, you got any glsa updates?"
21:31 <+wolf31o2> when replicating a box, the replication client itself knows it needs a new ID
21:31 <@agaffney> wolf31o2: well, the unique ID could be based off the MAC address
21:31 <@agaffney> wolf31o2: so that you don't have to generate one
21:32 <+wolf31o2> it could
21:32 <@blackace> codeman: and the server would check it's DB and see if the admin wanted that box to get automatic GLSA updates or not
21:32 <@agaffney> even if the MAC changes, that unique ID would be recorded somewhere
21:32 <+wolf31o2> blackace: that is actually exactly what satellite does... has a check box "This server gets automatic errata"
21:32 <@agaffney> wolf31o2: right now, glimd (or whatever) identifies clients my MAC
21:32 <@codeman> blackace: i agree.  server needs to be in charge of it all.. i was just throwing out the idea of client-control.
21:32 <@agaffney> s:my:by:
21:33 <+wolf31o2> agaffney: that's fine, so long as it is stored on the client... so even if the MAC changes the ID doesn't (until a new one is generated on the box)
21:33 <@agaffney> wolf31o2: so it wouldn't take much to have that not necessarily be MAC
21:33 <@agaffney> wolf31o2: but not have to change the client's code
21:33 <@agaffney> wolf31o2: easy enough
21:33 <@blackace> codeman: no no no, the server isn't in control from the point of view of linux permissions...ie. the server doesn't just tell the client to do stuff...the client has to ask "is there anything you want me to do?"
21:33 <+wolf31o2> right...
21:34 <@blackace> so if you have root on a client...you can turn off it's ask the server behavior
21:34 <@agaffney> codeman: the server may have the master list of stuff to do, but it can't force a client to do something
21:34 <@codeman> gotcha
21:34 <+wolf31o2> the client would be the authorative source... so like... the server could have "process errata automatically" enabled, and it is diabled in the client's config... well... it won't do it automatically
21:34 <@agaffney> codeman: the server doesn't tell the client what to do...the client asks the server what it wants the client to do
21:34 <+wolf31o2> exactly
21:35 <@codeman> alright.
21:35 <@codeman> so...
21:35 <+wolf31o2> it's all like... "yo server... what you got fo me?" and the server is like "I gotz DEEZ NUTZ!!!" and the client is like "step off, brother... I don't need your jive"
21:36  * agaffney falls over
21:36 <@codeman> i see groups of updates, rollback, ... what else is on our shopping cart list?
21:36 <@agaffney> codeman: I want some ice cream!
21:36 <+wolf31o2> personally, I would like to see (eventually) monitoring
21:36 <@codeman> hrm
21:37 <+wolf31o2> maybe some integration with nagios/cacti/whatever (don't really care)
21:37 <@codeman> i talked with people about that at work, they all said they'd prefer to just let their current monitoring software (nagios) do its job
21:37 <@codeman> but i think we can make nagios just another module to GLMD
21:37 <+wolf31o2> well... you could use nagios with USE="noweb" and provide your own interface
21:37 <+wolf31o2> right
21:38 <+wolf31o2> so long as you talk to the socket right, nagios doesn't give a damn what the actual web interface is
21:38 <@codeman> more like provide nagios with info it wants
21:38 <+wolf31o2> they're decoupled very well
21:38 <@codeman> well i don't see the point in rewriting their web interface
21:38 <@codeman> it's quite nice
21:38  * wolf31o2 knows all about writing nagios provisioning
21:38 <@codeman> kinda, if you theme it
21:38 <+wolf31o2> that would be fine too... for now
21:38 <+wolf31o2> heh
21:39 <+wolf31o2> I know I would end up writing a new front-end to match what we have
21:39 <@codeman> what kindof info does nagios need?
21:39 <+wolf31o2> give a universal interface
21:39 <+wolf31o2> from the web front-end?
21:39 <+wolf31o2> nothing
21:39 <@codeman> eh.. no
21:39  * codeman isn't phrasing this well
21:39 <+wolf31o2> you mean wrt what to check, etc?
21:39 <@codeman> like its a daemon that runs on each machine, right
21:39 <+wolf31o2> nope
21:40 <+wolf31o2> it runs on the server and that's it
21:40 <@codeman> so how does it know which machines to check?
21:40 <+wolf31o2> now... it *can* have a daemon on the machines (if you need it)
21:40 <+wolf31o2> config files on the server
21:40 <+wolf31o2> like... my nagios config is all snmp
21:40 <+wolf31o2> well... snmp and actual service checks
21:40 <@codeman> so perhaps we could use our client information stored in GLMD and give all that info over to nagios
21:41 <+wolf31o2> like... it checks my web site/postfix/imap by actually connecting and checking for valid responses
21:41 <+wolf31o2> exactly
21:41 <@codeman> so we could tell it which machines there are
21:41 <@codeman> that would rock
21:41 <+wolf31o2> hell... I already have code to do something like that... it would need very minimal adjustment...
21:45 <@codeman> what steps do we need to take to make ourselves a new official gentoo project?
21:45 <+wolf31o2> codeman: for this?
21:46 <@codeman> for the GLMD stuff
21:46 <@agaffney> let's wait and see where it goes, first
21:46 <+wolf31o2> codeman: as far as I know... just setup a project page...
21:46 <+wolf31o2> yeah
21:46 <@codeman> i'm not quite sure i see it being needing more than one project
21:46 <+wolf31o2> me either
21:47 <@codeman> i mean the stuff we're talking about is complicated and all, but still quite doable
21:47 <@codeman> but i'm quite certain there's a bunch more devs that'd be interested in this thing
21:47 <@agaffney> are you talking starting a new higher-level project and pulling the installer under that?
21:47 <@codeman> since a lot of devs are sys admins
21:47 <+wolf31o2> personally, I'd rather we start working on some stuff and get something out there... even if just mockups, etc
21:48 <@codeman> agaffney: eh, not necessarially
21:48 <@agaffney> wolf31o2: same here
21:48 <+wolf31o2> before unleashing it on the multi-directional tug
21:48 <+wolf31o2> because I think the moment people get interested, it ends up GLEP19
21:48 <+wolf31o2> 50 people wanting it to go in 50 directions... and it goes nowhere
21:48 <@agaffney> heh
21:48 <@codeman> well my attempt with dreamweaver was a very quick thing.. don't consider it any real mockup
21:49 <@agaffney> a single page with a few links isn't exactly a mockup :)
21:49 <+wolf31o2> hehe
21:49 <+wolf31o2> I cna screenshot a bunch of satellite stuff tomorrow
21:49  * blackace would be happy to work on this in earnest...now we're getting into the stuff I like :)
21:49 <@codeman> i know.  it's pathetic, but i at least wanted to get a little of my vision out there
21:49 <+wolf31o2> to give you and idea of kinda what I mean
21:50 <@codeman> blackace: i know.  GLI got pulled towards the user end for a very long time.. but i think it all has ended up working for the best
21:50 <@codeman> now we've got a semi-stable installer
21:50 <@codeman> and that's a great base for everything else from now on
21:50 <@blackace> yeah, which is why I was content to just stfu :)
21:51 <@blackace> ...ok...stfu, _some_ of the time ;)
21:51 <+wolf31o2> heh
1:52  * codeman stops logging

------=_Part_6866_21781586.1134003167647--
-- 
[email protected] mailing list