announcing synfl, the next generation of fontlock
Marcus Harnisch <[email protected]> Tue, 11 May 2004 23:11:27 -0400
| Newsgroups | gmane.emacs.xemacs.design,gmane.emacs.semantic |
|---|---|
| Message-ID | <[email protected]> |
Hi Stephen, Great work! Any chance to get in touch with the CEDET folks (http://cedet.sourceforge.net/)? It feels kinda weird to have two different parsers run over the same file. I didn't spend much time looking at the code, but could both parsers maybe share stuff? The fontification backend could also be made modular to access either `semantic' or `synfl' data. Sharing the same grammar source would already be a great thing to avoid having to maintain two grammars for each language. In case you were going to ask -- I am not volunteering ;-) Best regards, Marcus Stephen J. Turnbull writes: > That should get your attention! > > Returning to reason and modesty, this is a rather primitive module > based on an LR parser engine and an SLR parser generator. The parser > wants tuning and optimization, but is perfectly servicable for > experimentation. The parser generator probably should be extended to > a LALR parser generator for general use, but that should be reasonably > straightforward (though a little time-consuming) for anybody with a > copy of The Dragon Book to hand. (Aho, Sethi, and Ullman, _Compilers: > ..._) > > I probably shouldn't play with it much for the next few weeks, so, now > that it's basically working, here it is: > > cvs -d :pserver:[email protected]:/pack/xemacscvs \ > checkout -d synfl XEmacs/packages/unsupported/stephen/synfl > > http://turnbull.sk.tsukuba.ac.jp/Tools/XEmacs/synfl.tar.gz > > To use it, you should just byte-compile everything and load up. > > The Makefile is untested. It expects to find XEmacs.rules in the > grandparent directory, as usual. "distribution" in package-info.in is > set to "unsupported", so "make install" is not a good idea; I don't > think it has a chance of working at all. > > Features: > > There's a built-in trace/debug facility, so the package is of > educational interest. You can watch the LR parser, and the SLR > parser-generator, as they do their things. I thought it was pretty > cool, but maybe that's just me. > > The token stream implementation used to catch the results of lexical > analysis is probably OK. The "interface" used to mesh the token > stream, the grammar, and the parser output is an abortion; cleaning > that up is the next task, so my apologies in advance to anybody who > spontaneously combusts from looking at it. It just didn't want to be > abstracted yet, but I tried to force it.... > > The font-lock capability looks to be about twice as fast as (uncached, > nonlazy) traditional font-lock already, but all of that is due to not > using complex regular expressions; the syntax analyzer is currently > just along for the ride, wasting a small amount of time in the > interest of making sure it's reasonably efficient. It is. The C > grammar used so far is trivial, just a stream of tokens. > > Profile data for fontlock+cc-mode, a lock-c grammar, and a lex'n'lock > implementation applied to src/emacs.c are in doc/profile_emacs.c_*. > Tick counts are purely notional, as the XEmacs used for profiling is > fully pessimized. (Sorry, Hrvoje. :-) I will say that since so much > of the time for font-lock is spent in regexp matching (all in C), and > so little is for synfl, I'll bet that the advantage for synfl > _increases_ with optimization of XEmacs. > > It probably wouldn't be too difficult to arrange to call the parser > recursively on a different grammar (for the same kinds of files that > mmm-mode is used for). Haven't done that, but I have given a little > thought to avoiding introducing code that would prevent it. > > The lex'n'lock strategy is of independent interest. What's giving the > speed for both the synfl locker and the lex'n'locker is to lex the > tokens and then figure out what the keywords are by looking up the > tokens in a hash table---this can be done without the parser stuff, > and it's at least 5 times as fast as traditional font-lock. I am not > yet willing to conjecture whether either strategy will handle bogus > syntax or re-font-lock on-the-fly and still maintain efficiency. > > Documentation, outside of the sources, is non-existent, of course. > They are pretty extensively commented, cross-referenced to the Dragon > Book, and the docstrings are fairly good, I hope. You know where to > find me if you have questions. :-) > > -- > Institute of Policy and Planning Sciences http://turnbull.sk.tsukuba.ac.jp > University of Tsukuba Tennodai 1-1-1 Tsukuba 305-8573 JAPAN > Ask not how you can "do" free software business; > ask what your business can "do for" free software. >