Re: Announcing C::Blocks, a different way to interface Perl and C code
[email protected] (Ingy dot Net) Fri, 23 May 2014 14:38:04 -0700
| Newsgroups | perl.xs,perl.inline |
|---|---|
| Message-ID | <CAHJtQJ4eyJ+O6EO6hb=p9wn6dxbyTPCWY6oKOgz63YxesDcTfg@mail.gmail.com> |
--089e0149ce266c7e8d04fa180be4 Content-Type: text/plain; charset=UTF-8 I for one welcome our new C::Blocks overlords! :-) On Fri, May 23, 2014 at 5:35 AM, David Mertens <[email protected]>wrote: > Hey everyone, > > tl;dr: C::Blocks is a new TinyCC-based module, presently only available on > github. (1) It jit-compiles blocks of C code, building and inserting OPs > into the Perl OP tree, making invocation of C code essentially free. (2) It > will allow different blocks of C code to share function and struct > declarations, thus removing the need to always recompile perl.h, an > otherwise major cost of jit-compiling C code that can interface with Perl > and Perl data structures. > > I am currently seeking help and encouragement to squash the segfaults that > currently prevent the completion of the second feature. :-) > > ---- > > I like Perl, but I like C, too. I would like to be able to write and call > C code from Perl in about as painless a way as possible. Inline::C is nice, > as are XS::TCC and C::TinyCompiler, but we can do better. C::Blocks is my > attempt to do better. > > *Pain point 1*. C code should be a first class citizen. With XS::TCC and > C::TinyCompiler, you pass your code to the compiler via a string. With > Inline::C, you either place your code at the bottom of your script in a > __DATA__ section, or you enclose it in a string. Steffen's module is > probably the most transparent in this sense. Still, working with an > interface that requires me to compile a string to get my product feels the > same as compiling a regex from a string. This is Perl! We can do better! > > C::Blocks does better by using a keyword parser hook. Blocks of code that > you want executed are called like so: > > print "Before cblock\n"; > cblock { > printf("In cblock\n"); > } > print "After cblock\n"; > > If stdio.h is included, you get the output > > Before cblock > In cblock > After cblock > > Because it uses a parser hook, the C code really is inline with your Perl > code. > > *Pain point 2*. Calling C code should be obvious and cheap. All three > modules discussed so far provide a mechanism for calling C functions. This > means that for a simple, small operation, I must wrap my idea into a > function one place and invoke it in another. Furthermore, if I want to > repeatedly call a block of C code in a loop, I must define that block of > code somewhere outside of the loop, potentially very far from the call > site. C::TinyCompiler suffers further because it uses a complicated and > rather slow calling mechanism. > > C::Blocks solves this by extracting and jit-compiling the C code at Perl > parse time, generating an OP and inserting it into the Perl OP tree. This > means that you can insert your C code exactly where you want it and not > worry about repeated re-compiles. If you were to wrap the example given > above in a for loop, you would see how this works. > > *Pain point 3*. Sharing C code should be as easy as sharing Perl code. > C::TinyCompiler provides a fairly complex mechanism to allow modules to add > declarations and symbols to a compiler context. Any string that uses that > will need to recompile those declarations, however, tempting me to > prematurely optimize by placing all of my C code in one giant string > instead of interspersed among my Perl code. Neither XS::TCC nor Inline::C > provide much (if any) automated machinery to share code. > > C::Blocks provides a mechanism to share function declarations, struct > definitions, and other identifiers with other cblocks in the current > lexical scope, as well as to share them on a per-package basis. It is even > more versatile than normal Perl function scoping, allowing you to correctly > correlate functionality with lexical scope. (It is also somewhat buggy, as > discussed next.) > > *Pain point 4*. Changing C code should not cost anything. Inline::C can > take seconds to recompile a changed set of C code. In contrast, there is no > cost associated with changing code when using XS::TCC and C::TinyCompiler > because they jit-compile their code. That comes at the cost, however, of > always compiling everything each time you invoke your Perl script. If your > C code needs the Perl C API, you will have to re-parse perl.h every time > you *compile* a code block, which can happen many times with each > execution of your script. Inline::C's caching mechanism provides a big win > in that respect, unless you change your code. It would be nice if we could > somehow cache the result of parsing ``#include "perl.h"''. > > C::Blocks uses a fork of tcc that I've been working on for many months > aimed at allowing one compiler context to share its symbol table with other > compiler contexts. This is related to the previous point. The sharing > mechanism discussed in the previous point applies to preprocessor includes, > so once I have compiled a block that uses the Perl headers, I can share all > of those declarations with later compilation units, without recompiling. In > future work, I plan to store these symbol tables to disk so that they don't > even need to be re-parsed each time you run your script. > > ---- > > C::Blocks currently addresses, completely, pain points 1 and 2 above. It > has taken many months to hack on tcc to reach this point. I have now > encountered some segfault-causing issues when trying to share code, yet I > have a hard time reproducing those segfaults with direct tests on tcc. If > you think this sounds like a cool project, I would appreciate some > camaraderie as I try to dig into the internals of tcc. You can find me on > perl's IRC network hanging out on #pdl, #xs, and #tinycc, among other > channels. You can find my work at https://github.com/run4flat/C-Blocks > > If you would like to help out, let me know and I will give you a tour > through the codebase. :-) > > Any help or encouragement would be much appreciated! Thanks! > David > > -- > "Debugging is twice as hard as writing the code in the first place. > Therefore, if you write the code as cleverly as possible, you are, > by definition, not smart enough to debug it." -- Brian Kernighan > --089e0149ce266c7e8d04fa180be4 Content-Type: text/html; charset=UTF-8 Content-Transfer-Encoding: quoted-printable <div dir=3D"ltr">I for one welcome our new C::Blocks overlords! :-)<br></di= v><div class=3D"gmail_extra"><br><br><div class=3D"gmail_quote">On Fri, May= 23, 2014 at 5:35 AM, David Mertens <span dir=3D"ltr"><<a href=3D"mailto= :[email protected]" target=3D"_blank">[email protected]</a>&g= t;</span> wrote:<br> <blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p= x #ccc solid;padding-left:1ex"><div dir=3D"ltr"><div><div><div><div>Hey eve= ryone,<br><br></div><div>tl;dr: C::Blocks is a new TinyCC-based module, pre= sently only available on github. (1) It jit-compiles blocks of C code, buil= ding and inserting OPs into the Perl OP tree, making invocation of C code e= ssentially free. (2) It will allow different blocks of C code to share func= tion and struct declarations, thus removing the need to always recompile pe= rl.h, an otherwise major cost of jit-compiling C code that can interface wi= th Perl and Perl data structures.<br> </div><div><br></div><div>I am currently seeking help and encouragement to = squash the segfaults that currently prevent the completion of the second fe= ature. :-)<br></div><div><br>----<br><br></div><div>I like Perl, but I like= C, too. I would like to be able to write and call C code from Perl in abou= t as painless a way as possible. Inline::C is nice, as are XS::TCC and C::T= inyCompiler, but we can do better. C::Blocks is my attempt to do better.<br= > <br><b>Pain point 1</b>. C code should be a first class citizen. With XS::T= CC and C::TinyCompiler, you pass your code to the compiler via a string. Wi= th Inline::C, you either place your code at the bottom of your script in a = __DATA__ section, or you enclose it in a string. Steffen's module is pr= obably the most transparent in this sense. Still, working with an interface= that requires me to compile a string to get my product feels the same as c= ompiling a regex from a string. This is Perl! We can do better!<br> <br></div><div>C::Blocks does better by using a keyword parser hook. Blocks= of code that you want executed are called like so:<br><br></div><div>=C2= =A0=C2=A0=C2=A0 print "Before cblock\n";<br></div><div>=C2=A0=C2= =A0=C2=A0 cblock {<br></div> =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 printf("In cblock\n");= <br> <div>=C2=A0=C2=A0=C2=A0 }<br></div><div>=C2=A0=C2=A0=C2=A0 print "Afte= r cblock\n";<br></div><div><br></div><div>If stdio.h is included, you = get the output<br><br></div><div>=C2=A0=C2=A0=C2=A0 Before cblock<br></div>= <div>=C2=A0=C2=A0=C2=A0 In cblock<br></div><div>=C2=A0=C2=A0=C2=A0 After cb= lock<br> </div><div><br></div><div>Because it uses a parser hook, the C code really = is inline with your Perl code.<br><br></div><div><b>Pain point 2</b>. Calli= ng C code should be obvious and cheap. All three modules discussed so far p= rovide a mechanism for calling C functions. This means that for a simple, s= mall operation, I must wrap my idea into a function one place and invoke it= in another. Furthermore, if I want to repeatedly call a block of C code in= a loop, I must define that block of code somewhere outside of the loop, po= tentially very far from the call site. C::TinyCompiler suffers further beca= use it uses a complicated and rather slow calling mechanism.<br> <br></div><div>C::Blocks solves this by extracting and jit-compiling the C = code at Perl parse time, generating an OP and inserting it into the Perl OP= tree. This means that you can insert your C code exactly where you want it= and not worry about repeated re-compiles. If you were to wrap the example = given above in a for loop, you would see how this works.<br> <br></div><div><b>Pain point 3</b>. Sharing C code should be as easy as sha= ring Perl code. C::TinyCompiler provides a fairly complex mechanism to allo= w modules to add declarations and symbols to a compiler context. Any string= that uses that will need to recompile those declarations, however, temptin= g me to prematurely optimize by placing all of my C code in one giant strin= g instead of interspersed among my Perl code. Neither XS::TCC nor Inline::C= provide much (if any) automated machinery to share code.<br> <br></div><div>C::Blocks provides a mechanism to share function declaration= s, struct definitions, and other identifiers with other cblocks in the curr= ent lexical scope, as well as to share them on a per-package basis. It is e= ven more versatile than normal Perl function scoping, allowing you to corre= ctly correlate functionality with lexical scope. (It is also somewhat buggy= , as discussed next.)<br> </div><div><br></div><div><b>Pain point 4</b>. Changing C code should not c= ost anything. Inline::C can take seconds to recompile a changed set of C co= de. In contrast, there is no cost associated with changing code when using = XS::TCC and C::TinyCompiler because they jit-compile their code. That comes= at the cost, however, of always compiling everything each time you invoke = your Perl script. If your C code needs the Perl C API, you will have to re-= parse perl.h every time you <i>compile</i> a code block, which can happen m= any times with each execution of your script. Inline::C's caching mecha= nism provides a big win in that respect, unless you change your code. It wo= uld be nice if we could somehow cache the result of parsing ``#include &quo= t;perl.h"''.<br> <br></div><div>C::Blocks uses a fork of tcc that I've been working on f= or many months aimed at allowing one compiler context to share its symbol t= able with other compiler contexts. This is related to the previous point. T= he sharing mechanism discussed in the previous point applies to preprocesso= r includes, so once I have compiled a block that uses the Perl headers, I c= an share all of those declarations with later compilation units, without re= compiling. In future work, I plan to store these symbol tables to disk so t= hat they don't even need to be re-parsed each time you run your script.= <br> </div><div><br></div>----<br><br></div><div>C::Blocks currently addresses, = completely, pain points 1 and 2 above. It has taken many months to hack on = tcc to reach this point. I have now encountered some segfault-causing issue= s when trying to share code, yet I have a hard time reproducing those segfa= ults with direct tests on tcc. If you think this sounds like a cool project= , I would appreciate some camaraderie as I try to dig into the internals of= tcc. You can find me on perl's IRC network hanging out on #pdl, #xs, a= nd #tinycc, among other channels. You can find my work at <a href=3D"https:= //github.com/run4flat/C-Blocks" target=3D"_blank">https://github.com/run4fl= at/C-Blocks</a><br> <br></div><div>If you would like to help out, let me know and I will give y= ou a tour through the codebase. :-)<br></div><div><br></div><div>Any help o= r encouragement would be much appreciated! Thanks!<span class=3D"HOEnZb"><f= ont color=3D"#888888"><br> David<br clear=3D"all"> </font></span></div></div></div><span class=3D"HOEnZb"><font color=3D"#8888= 88"><div><div><div><div><div><br>-- <br>=C2=A0"Debugging is twice as h= ard as writing the code in the first place.<br> =C2=A0 Therefore, if you write the code as cleverly as possible, you are,<= br> =C2=A0 by definition, not smart enough to debug it." -- Brian Kernigh= an<br> </div></div></div></div></div></font></span></div> </blockquote></div><br></div> --089e0149ce266c7e8d04fa180be4--