Re: Announcing C::Blocks, a different way to interface Perl and C code

[email protected] (Ingy dot Net) Fri, 23 May 2014 14:38:04 -0700
Newsgroups perl.xs,perl.inline
Message-ID <CAHJtQJ4eyJ+O6EO6hb=p9wn6dxbyTPCWY6oKOgz63YxesDcTfg@mail.gmail.com>
--089e0149ce266c7e8d04fa180be4
Content-Type: text/plain; charset=UTF-8

I for one welcome our new C::Blocks overlords! :-)


On Fri, May 23, 2014 at 5:35 AM, David Mertens <[email protected]>wrote:

> Hey everyone,
>
> tl;dr: C::Blocks is a new TinyCC-based module, presently only available on
> github. (1) It jit-compiles blocks of C code, building and inserting OPs
> into the Perl OP tree, making invocation of C code essentially free. (2) It
> will allow different blocks of C code to share function and struct
> declarations, thus removing the need to always recompile perl.h, an
> otherwise major cost of jit-compiling C code that can interface with Perl
> and Perl data structures.
>
> I am currently seeking help and encouragement to squash the segfaults that
> currently prevent the completion of the second feature. :-)
>
> ----
>
> I like Perl, but I like C, too. I would like to be able to write and call
> C code from Perl in about as painless a way as possible. Inline::C is nice,
> as are XS::TCC and C::TinyCompiler, but we can do better. C::Blocks is my
> attempt to do better.
>
> *Pain point 1*. C code should be a first class citizen. With XS::TCC and
> C::TinyCompiler, you pass your code to the compiler via a string. With
> Inline::C, you either place your code at the bottom of your script in a
> __DATA__ section, or you enclose it in a string. Steffen's module is
> probably the most transparent in this sense. Still, working with an
> interface that requires me to compile a string to get my product feels the
> same as compiling a regex from a string. This is Perl! We can do better!
>
> C::Blocks does better by using a keyword parser hook. Blocks of code that
> you want executed are called like so:
>
>     print "Before cblock\n";
>     cblock {
>         printf("In cblock\n");
>     }
>     print "After cblock\n";
>
> If stdio.h is included, you get the output
>
>     Before cblock
>     In cblock
>     After cblock
>
> Because it uses a parser hook, the C code really is inline with your Perl
> code.
>
> *Pain point 2*. Calling C code should be obvious and cheap. All three
> modules discussed so far provide a mechanism for calling C functions. This
> means that for a simple, small operation, I must wrap my idea into a
> function one place and invoke it in another. Furthermore, if I want to
> repeatedly call a block of C code in a loop, I must define that block of
> code somewhere outside of the loop, potentially very far from the call
> site. C::TinyCompiler suffers further because it uses a complicated and
> rather slow calling mechanism.
>
> C::Blocks solves this by extracting and jit-compiling the C code at Perl
> parse time, generating an OP and inserting it into the Perl OP tree. This
> means that you can insert your C code exactly where you want it and not
> worry about repeated re-compiles. If you were to wrap the example given
> above in a for loop, you would see how this works.
>
> *Pain point 3*. Sharing C code should be as easy as sharing Perl code.
> C::TinyCompiler provides a fairly complex mechanism to allow modules to add
> declarations and symbols to a compiler context. Any string that uses that
> will need to recompile those declarations, however, tempting me to
> prematurely optimize by placing all of my C code in one giant string
> instead of interspersed among my Perl code. Neither XS::TCC nor Inline::C
> provide much (if any) automated machinery to share code.
>
> C::Blocks provides a mechanism to share function declarations, struct
> definitions, and other identifiers with other cblocks in the current
> lexical scope, as well as to share them on a per-package basis. It is even
> more versatile than normal Perl function scoping, allowing you to correctly
> correlate functionality with lexical scope. (It is also somewhat buggy, as
> discussed next.)
>
> *Pain point 4*. Changing C code should not cost anything. Inline::C can
> take seconds to recompile a changed set of C code. In contrast, there is no
> cost associated with changing code when using XS::TCC and C::TinyCompiler
> because they jit-compile their code. That comes at the cost, however, of
> always compiling everything each time you invoke your Perl script. If your
> C code needs the Perl C API, you will have to re-parse perl.h every time
> you *compile* a code block, which can happen many times with each
> execution of your script. Inline::C's caching mechanism provides a big win
> in that respect, unless you change your code. It would be nice if we could
> somehow cache the result of parsing ``#include "perl.h"''.
>
> C::Blocks uses a fork of tcc that I've been working on for many months
> aimed at allowing one compiler context to share its symbol table with other
> compiler contexts. This is related to the previous point. The sharing
> mechanism discussed in the previous point applies to preprocessor includes,
> so once I have compiled a block that uses the Perl headers, I can share all
> of those declarations with later compilation units, without recompiling. In
> future work, I plan to store these symbol tables to disk so that they don't
> even need to be re-parsed each time you run your script.
>
> ----
>
> C::Blocks currently addresses, completely, pain points 1 and 2 above. It
> has taken many months to hack on tcc to reach this point. I have now
> encountered some segfault-causing issues when trying to share code, yet I
> have a hard time reproducing those segfaults with direct tests on tcc. If
> you think this sounds like a cool project, I would appreciate some
> camaraderie as I try to dig into the internals of tcc. You can find me on
> perl's IRC network hanging out on #pdl, #xs, and #tinycc, among other
> channels. You can find my work at https://github.com/run4flat/C-Blocks
>
> If you would like to help out, let me know and I will give you a tour
> through the codebase. :-)
>
> Any help or encouragement would be much appreciated! Thanks!
> David
>
> --
>  "Debugging is twice as hard as writing the code in the first place.
>   Therefore, if you write the code as cleverly as possible, you are,
>   by definition, not smart enough to debug it." -- Brian Kernighan
>

--089e0149ce266c7e8d04fa180be4
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr">I for one welcome our new C::Blocks overlords! :-)<br></di=
v><div class=3D"gmail_extra"><br><br><div class=3D"gmail_quote">On Fri, May=
 23, 2014 at 5:35 AM, David Mertens <span dir=3D"ltr">&lt;<a href=3D"mailto=
:[email protected]" target=3D"_blank">[email protected]</a>&g=
t;</span> wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex"><div dir=3D"ltr"><div><div><div><div>Hey eve=
ryone,<br><br></div><div>tl;dr: C::Blocks is a new TinyCC-based module, pre=
sently only available on github. (1) It jit-compiles blocks of C code, buil=
ding and inserting OPs into the Perl OP tree, making invocation of C code e=
ssentially free. (2) It will allow different blocks of C code to share func=
tion and struct declarations, thus removing the need to always recompile pe=
rl.h, an otherwise major cost of jit-compiling C code that can interface wi=
th Perl and Perl data structures.<br>

</div><div><br></div><div>I am currently seeking help and encouragement to =
squash the segfaults that currently prevent the completion of the second fe=
ature. :-)<br></div><div><br>----<br><br></div><div>I like Perl, but I like=
 C, too. I would like to be able to write and call C code from Perl in abou=
t as painless a way as possible. Inline::C is nice, as are XS::TCC and C::T=
inyCompiler, but we can do better. C::Blocks is my attempt to do better.<br=
>

<br><b>Pain point 1</b>. C code should be a first class citizen. With XS::T=
CC and C::TinyCompiler, you pass your code to the compiler via a string. Wi=
th Inline::C, you either place your code at the bottom of your script in a =
__DATA__ section, or you enclose it in a string. Steffen&#39;s module is pr=
obably the most transparent in this sense. Still, working with an interface=
 that requires me to compile a string to get my product feels the same as c=
ompiling a regex from a string. This is Perl! We can do better!<br>

<br></div><div>C::Blocks does better by using a keyword parser hook. Blocks=
 of code that you want executed are called like so:<br><br></div><div>=C2=
=A0=C2=A0=C2=A0 print &quot;Before cblock\n&quot;;<br></div><div>=C2=A0=C2=
=A0=C2=A0 cblock {<br></div>
=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 printf(&quot;In cblock\n&quot;);=
<br>
<div>=C2=A0=C2=A0=C2=A0 }<br></div><div>=C2=A0=C2=A0=C2=A0 print &quot;Afte=
r cblock\n&quot;;<br></div><div><br></div><div>If stdio.h is included, you =
get the output<br><br></div><div>=C2=A0=C2=A0=C2=A0 Before cblock<br></div>=
<div>=C2=A0=C2=A0=C2=A0 In cblock<br></div><div>=C2=A0=C2=A0=C2=A0 After cb=
lock<br>

</div><div><br></div><div>Because it uses a parser hook, the C code really =
is inline with your Perl code.<br><br></div><div><b>Pain point 2</b>. Calli=
ng C code should be obvious and cheap. All three modules discussed so far p=
rovide a mechanism for calling C functions. This means that for a simple, s=
mall operation, I must wrap my idea into a function one place and invoke it=
 in another. Furthermore, if I want to repeatedly call a block of C code in=
 a loop, I must define that block of code somewhere outside of the loop, po=
tentially very far from the call site. C::TinyCompiler suffers further beca=
use it uses a complicated and rather slow calling mechanism.<br>

<br></div><div>C::Blocks solves this by extracting and jit-compiling the C =
code at Perl parse time, generating an OP and inserting it into the Perl OP=
 tree. This means that you can insert your C code exactly where you want it=
 and not worry about repeated re-compiles. If you were to wrap the example =
given above in a for loop, you would see how this works.<br>

<br></div><div><b>Pain point 3</b>. Sharing C code should be as easy as sha=
ring Perl code. C::TinyCompiler provides a fairly complex mechanism to allo=
w modules to add declarations and symbols to a compiler context. Any string=
 that uses that will need to recompile those declarations, however, temptin=
g me to prematurely optimize by placing all of my C code in one giant strin=
g instead of interspersed among my Perl code. Neither XS::TCC nor Inline::C=
 provide much (if any) automated machinery to share code.<br>

<br></div><div>C::Blocks provides a mechanism to share function declaration=
s, struct definitions, and other identifiers with other cblocks in the curr=
ent lexical scope, as well as to share them on a per-package basis. It is e=
ven more versatile than normal Perl function scoping, allowing you to corre=
ctly correlate functionality with lexical scope. (It is also somewhat buggy=
, as discussed next.)<br>

</div><div><br></div><div><b>Pain point 4</b>. Changing C code should not c=
ost anything. Inline::C can take seconds to recompile a changed set of C co=
de. In contrast, there is no cost associated with changing code when using =
XS::TCC and C::TinyCompiler because they jit-compile their code. That comes=
 at the cost, however, of always compiling everything each time you invoke =
your Perl script. If your C code needs the Perl C API, you will have to re-=
parse perl.h every time you <i>compile</i> a code block, which can happen m=
any times with each execution of your script. Inline::C&#39;s caching mecha=
nism provides a big win in that respect, unless you change your code. It wo=
uld be nice if we could somehow cache the result of parsing ``#include &quo=
t;perl.h&quot;&#39;&#39;.<br>

<br></div><div>C::Blocks uses a fork of tcc that I&#39;ve been working on f=
or many months aimed at allowing one compiler context to share its symbol t=
able with other compiler contexts. This is related to the previous point. T=
he sharing mechanism discussed in the previous point applies to preprocesso=
r includes, so once I have compiled a block that uses the Perl headers, I c=
an share all of those declarations with later compilation units, without re=
compiling. In future work, I plan to store these symbol tables to disk so t=
hat they don&#39;t even need to be re-parsed each time you run your script.=
<br>

</div><div><br></div>----<br><br></div><div>C::Blocks currently addresses, =
completely, pain points 1 and 2 above. It has taken many months to hack on =
tcc to reach this point. I have now encountered some segfault-causing issue=
s when trying to share code, yet I have a hard time reproducing those segfa=
ults with direct tests on tcc. If you think this sounds like a cool project=
, I would appreciate some camaraderie as I try to dig into the internals of=
 tcc. You can find me on perl&#39;s IRC network hanging out on #pdl, #xs, a=
nd #tinycc, among other channels. You can find my work at <a href=3D"https:=
//github.com/run4flat/C-Blocks" target=3D"_blank">https://github.com/run4fl=
at/C-Blocks</a><br>

<br></div><div>If you would like to help out, let me know and I will give y=
ou a tour through the codebase. :-)<br></div><div><br></div><div>Any help o=
r encouragement would be much appreciated! Thanks!<span class=3D"HOEnZb"><f=
ont color=3D"#888888"><br>
David<br clear=3D"all">
</font></span></div></div></div><span class=3D"HOEnZb"><font color=3D"#8888=
88"><div><div><div><div><div><br>-- <br>=C2=A0&quot;Debugging is twice as h=
ard as writing the code in the first place.<br>
 =C2=A0 Therefore, if you write the code as cleverly as possible, you are,<=
br>
 =C2=A0 by definition, not smart enough to debug it.&quot; -- Brian Kernigh=
an<br>
</div></div></div></div></div></font></span></div>
</blockquote></div><br></div>

--089e0149ce266c7e8d04fa180be4--