Zencastr
00:00:00
00:00:01
Speed1x
Format▸
Share
Embed
Report

How Would You Type-Check 30 Million Lines of PHP? (with Julien Verlaguet)

Developer Voices
Developer Voices

1 plays · Sep 25, 2026

Transcript

Speaker: How would you manage 30 million lines of PHP? I know plenty of people would say, I would run away from that job screaming. But let's be fair, right? This isn't a PHP question. This is a 30 million lines of code question.

Speaker: And 30 million lines of anything is terrifying. How would you do it? You don't get to rewrite it. You can't say I'll rewrite it in my favorite language because rewriting that much code is such a long project. You'd be fired long before you delivered.

Speaker: You have to play the code base as it's dealt. Well, my guest this week was at Facebook when they faced this question back in about 2011. And the answer they came up with was add a type system.

Speaker: Add a type system to PHP and that's how the language hack was born. But as the code designer, Julien Valaguet, found out, adding a type system to PHP isn't actually the hardest part.

Speaker: The harder part is building a type system that people are happy to use. Because think about it. You've got all these developers who are used to making a couple of lines of code change, pressing refresh, and seeing the results.

Speaker: You can't suddenly tell them they've got to wait occasionally while we recompile two million lines of code in the dependency tree. So the real question is, how do you write tools and languages that can incrementally manage 30 million lines of code?

Speaker: And the answer to that question is a fascinating look at the realities of building a production compiler. It's also an approach to language and tooling and high performance incremental design that Julian's continued exploring ever since, for the past 15 years, all the way to his latest language, Skip.

Speaker: This is a fascinating one to record, especially if you're like me and you like languages and real-time systems. So let's get into it. I'm your host, Chris Jenkins. This is Developer Voices. And today's voice is Julian Villeguette.

Speaker: Joining me today is Julien Verlaget. Julien, are you doing? I'm good. How are you, Chris? Good to hear I'm very well. I'm very well. you are You're traveling at the moment, I gather. You're normally in Barcelona and you're in Paris. I'm in Paris right now, yeah. And it's I picked up one time a year where it's actually hot in Paris.

Speaker: And it's incredibly painful there because there isn't much air conditioning anywhere. But I managed to found a room to find a room with air conditioning. So I'm i'm surviving. That's absolutely vital at the moment, even in rainy old England that's vital.

Speaker: But um so I got you in to talk about reactive systems and about Skiplang and the Skiplang reactive framework. But then when we got chatting earlier,

Speaker: The backstory to this is so interesting. I think we're going to take the long and scenic route to get there. All right. take me back to the foundations. 2010, you're working at Facebook, right? Yeah. So I started, i think my first day was February 2011, if I remember well. But yeah, it's around that time. And yes, my first job at Facebook was to build um static analysis to find security holes. So that was really the the first job. so this was ah And this is static analysis of PHP files? Yeah, so this was an abstract interpreter. And the way it was working is it was basically doing a basic taint analysis, trying to get um strings that were not properly escaped

Speaker: and see sorry and see if they were getting into dangerous things, such as a sql injection SQL query, which would be a SQL injection, obviously a shell injection, but also XSS attacks and all sorts of different security attacks. and it was This is like when the user supplies a user supply value from the outside world. You track all the places it gets used to see if it ends up in an SQL query. Yeah, so yeah it could be yeah it could be from a database, it could be from a log, it could be from directly provided by the user, it could be anything that needs escaping and and needs to be tracked, basically. And so right it was working relatively well because back then,

Speaker: the there was still a bunch of code that was old code that was still written with an old school PHP style. And the old school PHP style was very much like, I don't know if any of you have written that before, but it was pretty much like HTML in a page, right? Like you would yeah you would basically have a top level, um you would basically do pretty much everything at top level, right?

Speaker: And yeah when that's the case, your your file is like a gigantic function. And when that's the case, we were actually able to track this without any types. We were able to analyze that because there was not enough going on that that you know we we could actually cover the whole thing. And then right what was happening is that the code base was modernizing at a very rapid pace.

Speaker: And so there was a lot of structure that was added and levels of indirection. And following those levels of indirections was was becoming harder and harder. And that's when the idea of um adding type information, static type information to PHP was floated around. And the first...

Speaker: version was not meant to be something that was going to be applied to the entire code base. It was meant to be used by security engineer or people who really care about correctness. you know Think about you know the people developing the crypto framework or stuff like that, to be able to use a stricter version of PHP, which back then we called strict mode.

Speaker: um and so Right. that's so the but but very quickly became clear that this wasn't going to really work because things were not very well isolated and so if you wanted to if you if you have type information that you can trust in one place but then the moment you leave that place you cannot trust what's going on anymore and so there was things were not as well split up as we would have hoped to.

Speaker: And so that's how we ended up developing a version of Hack that would work on the entire code base. Right. take Take me through that a little bit slower. why Why, when you're type checking one file or one place, would you not be able to trust it when you go to another place?

Speaker: So... um Imagine you have a file and this file is calling into a function um and that function cannot be trusted.

Speaker: So it comes back with, let's say it has declared that it comes back with a string, but in fact, this string can be null, right? Because it comes from a place that cannot be trusted. Now, the fact that this thing is potentially null,

Speaker: is basically part of your invariance without you realizing it, and it can creep in in all sorts of different places in the code base. And so it's difficult to build something trustworthy if you don't have... um if you don't have all the dependencies, so if you don't go bottom up.

Speaker: And going bottom up was very difficult because the dependencies of Facebook were a spaghetti bowl of dependencies. And so... I can believe that.

Speaker: So the the reason why we got there is basically there is this feature in PHP called autoloading. And so it's a nice feature when you're writing code, but basically what this feature does is whenever you're using a class or a function, if this class of function has never been seen before, um what will happen is the the VM is going to load it automatically for you, right? So imagine if this was Java,

Speaker: you're writing a class name and you don't have to do anything. The JVM is going to load the right jar for you, basically. see quote There's no importing, no explicit import. Exactly.

Speaker: The imports are deduced directly from ah from the names. um And so... I mean, it's nice because it's convenient, but what ended up happening is that there was millions and millions of lines of code that were written by thousands of engineers who never, ever thought about dependencies and never thought about you know how to split things up into smaller pieces.

Speaker: Right. Yeah. Yeah, I can believe that. It just magically works, so you magically hope it's well organized. Exactly. And what ends up happening is it's not well organized. So you end up with, you know, whenever you pull in a dependency, even for the most mundane thing, if you were writing, you know, 10 lines of of PHP, you would pull in millions of lines of dependencies. it It was always huge what you would pull in. And so that's when, you know, we decided for Hack to build a a language server. And so

Speaker: One of the things that, so we wanted we knew we wanted to build a statically typed version of PHP. and Right. we knew Wanted or had to? i Had to. It's true that if I could have started from scratch, I would have started from scratch for sure. Because there was a lot of baggage in in the language that I would have rather not have to deal with. But it was what it was. So, you know, it's what was there.

Speaker: So now if the idea is we're going to build a statically typed version of PHP, one of the main problems we are facing is the workflow of engineers at Facebook, the way they were working back then.

Speaker: So almost all the logic was in PHP back then. There was a few people writing you know some C++ plus plus on a corner, but most of the code, more than 70% of the code was PHP. And then we're talking tens of millions of lines here, right? Exactly. yeah And the the way it was working was very much like... um a website of the 90s, where very little client-side logic. You press a button that's that gets gets gets that issues a request to the the PHP server, and then basically all the business logic is server-side.

Speaker: um yeah And so these people, the way they would write code and the way they would iterate would be with, on the right side, ah a web browser, on the left side, the code editor, and they would just write code, press enter, refresh the browser and see if it works. right yeah Which is reasonable, right especially if you're trying to build a competing user experience.

Speaker: And so that feedback loop was very, very precious to them. And so there had been attempts in the organization to add linters, to add all sorts of things that were taking time, and people would just disable them.

Speaker: ah So if anything was taking too much time, they would just... ah get rid of the tool so that they could keep on iterating on their code. Yeah, right. So any compiler that's going to be successful or any static analysis tool that's going to be successful has to be incremental. Yeah. So the idea is if you want to insert yourself in that workflow, you have to be extremely fast. um And so, of course, our first instinct, as I as i was mentioning, ah i mean, the first thing thing when you want to make a language fast, is modularity, right? Split spit it up into smaller pieces. But that wasn't an option for us because of what i explained, that that there was a spaghetti bowl of dependencies and you really could not split things up into smaller libraries. So there was an effort within the company to do that. Somebody was trying to actually, you know,

Speaker: ah bring some sanity in in the in the landscape of dependencies. But he wasn't very successful and it wasn't really followed by the organization. And so I think it it had this project been successful, then we would have du definitely gone a more traditional route, probably like a type system, a type checker with different libraries and different modules. But because this was not an option,

Speaker: we ended up writing a language server. And so the challenge there was to have a language server over tens of millions of lines of code that actually responds very fast.

Speaker: um And so to do that, well, the the first type checker was written in OCaml, but back then OCaml didn't have um threading. um Now it does. It has multi-core OCaml.

Speaker: um And so it was a forked-based solution that was the very first implementation. ah But then the forks, what was dominating was the time to basically copy data around right between the different processes doing work. And so what we ended up doing in the end was a new runtime for CAMEL where there was forked processes, but then they share data through a shared heap.

Speaker: And this shared heap is can be accessed um concurrently, very efficiently. The one... um The one big constraint in this shared heap is that the the data put inside the shared heap are provably immutable.

Speaker: So that's that's the the one big constraint. But you can basically think of it this way. Imagine a gigantic atomic hash table where the keys and the values um have to be immutable.

Speaker: And then a bunch of processes, so Unix processes, not threads, that um orchestrate work with each other with one that controls the others to send tasks and do do work. And the orchestration is done through pipes, but the actual data is never sent through these pipes. They they are stored in this shared heap, in this atomic hash table. Okay, so you're using pipes for coordination, but this shared atomic hash map for sharing information.

Speaker: Before we go on, what's the what's the unit of work? Is it like one worker per file that's being type-checked? So it really depends on the phase. The first phase, which is parsing, yes, the unity of work is going to be file. And then after, in the subsequent phases, um it becomes much more fine-grained. It gets down to method function.

Speaker: um But um yeah, if you want, we can go through the first phase together. So um you yeah need to parse the files, right? So the first thing that we added was a daemon that was watching all the changes on all the files using iNotify.

Speaker: And the idea we wanted the daemon to, we wanted the the language server to keep up with what was going on on disk without having to wait for the user to hit refresh or sync or whatnot. So first, there's a daemon that watches basically all all the directories that of interest.

Speaker: Now, when you initialize the program, You start with, let's say, 10,000 files, right? So what you're going to do is, well, it's probably a lot more than that, but it doesn't matter. You have 10,000 files. So the first thing you do is you split them into buckets.

Speaker: And then say, I'm going to take buckets of 10, 20 files. I'm going to send those buckets to the workers that are going to do the work of parsing. So they take the files. They actually parse them.

Speaker: And then they put the results in shared memory. So now for each file name, I have associated to this file name an AST that lives in memory, that lives in shared memory. And at this stage, you're not the individual types, you're just the AST for the whole file? At this point, it's just the AST for the whole file. Right, okay. um So you finish this phase.

Speaker: Once this phase is finished, you have, for each file, you can look up an AST very quickly. Because the amount of data that we're going to put on in the shared heap is going to be humongous, what we did is we actually serialized them and compressed them.

Speaker: So they're both serialized and compressed. And we added caches for each worker. So whenever the worker would access the shared heap, it would not pay the cost of deservilization every single time if it was using the same data multiple times.

Speaker: Right, yeah. um So now we have finished phase one. We have all of our ASTs. Phase two, what we do is we declare all the types. So what that means is walking through all the classes and ah basically flattening all the the class hierarchies and all the interfaces so that for each class, we are able to quickly look up a method that um And so, for example, it also can mean specializing some generics. So let's say you have a class that extends a generic class, then we we will want to solve all these names, basically. So that in the type checking phase, whenever we want to write new a we want to be able to very quickly build an object that represents this class, the this object. Okay, so you're saying if you've got defined a generic list of E,

Speaker: um I'm using Java notation there. List of E, you are going to store in your big atomic hash map a key of list of string and list of integer for all the different concrete implementations you want. No, no, no. I explained that poorly. What we want is...

Speaker: Basically, for every single function, we want to be able to look up a signature. And for every single class, we want to be able to look at the class, but in a in a in a form where all the name have been solved. So you don't want, during type checking, to have to go up a class hierarchy to figure out what where this method is implemented. And if it is implemented, apply a bunch of substitution to get the type. right So imagine... Oh, OK.

Speaker: So that's that's just what I meant. So basically, at the end of this phase, you want a typing environment where all the types have been resolved. for classes and functions and constants. Right, so you're inlining all the definitions so that when you look up a specific thing, you get all of its children

Speaker: pre-dehydrated. Yeah, exactly. So you get all the children. ah Well, you get all the parents, not all the children. You get all the parents and you get them ah pre-substituted. So if if some of them were generic, then you solve the... Okay, right. I'm with you. All right. So now you have this and typing environment and now you're actually doing the real type checking.

Speaker: And so the real type checking, you same thing, you split up your files into buckets. You send those to the different to the different workers. And now the workers have all the pieces to do a good job. like They can first go look up the AST. So I cheated a little bit because there's an actual extra phase, but that doesn't bring that much two things. And so to simplify, there's a naming phase. that So you're not actually directly working on an AST, but it doesn't matter. So now the worker can look up the AST and it can walk the different um the different bodies of the different functions and different methods and all the code. And it has a typing environment and to look things up very quickly and very efficiently.

Speaker: um Right. So that's the language, so that's the initialization of the language server. And, um, That phase was actually really efficient. We were able to type check, I think it was um a million lines of code per second.

Speaker: so Really? Yeah, a very beefy machines. We had ah machines with 32 cores, if I remember well. That was the okay the the typical dev box back then was 32 cores.

Speaker: um That's not guaranteed though, because just having 32 cores doesn't guarantee you're making good use of them. So actually it was 32 processors and I think 64 cores. yeah And out of the 64 cores, we only used 56 at 100%.

Speaker: And the reason is the the bottleneck was the RAM, which is pretty much always the case whenever you're you're writing, whenever you're dealing with symbolic computation at scale, the the limiting factor is always um memory pressure. It's not so much um CPU itself.

Speaker: Yeah, I can believe that. yeah But we we still had really good performance because the RAM that we had was very performance and we had good throughput. But if if you're going to set up hardware for this kind of stuff, my advice is think very hard about memory. And that's probably going to be your bottleneck, not not your CPU, if you're working on static analysis and these kind of things.

Speaker: um So now we've walked through the... Initialization. So how do you think make that thing incremental? So yeah the way you make that thing incremental is basically every time something accesses the environment, accesses the heap, sorry.

Speaker: You have to keep track of those dependencies. So basically, if while you are type checking the the the function foo, you access type A, this needs to be recorded somewhere. You need to know the fact that foo depends on a And in fact, what interests you is the link the other way around. So you want, if A changes, you want to be able to know very quickly that you need to recheck foo, right?

Speaker: Yes, walking in the tree in one direction means you want to take a record, you want to record the fact that you're going to have to be notified in the other direction later. Exactly. Yeah, exactly. So we ended up building, handcrafting a very efficient data structure and very low level stuff with atomics and whatnot to make that really efficient because um if you the the dependencies have to be shared across all the different workers, because if they're not shared, then you have a synchronization process that is going to be horrible and suck up a lot of memory. I know because that's how we started. And then we ended up having to build this really efficient ah data structure to store dependencies ah very quickly.

Speaker: um And so the way it works in a nutshell is... You take um the name of what you want to, ah the name of what's on the left, the name of what's on the right, you um hash them, then you keep a 32-bit hash for both sides. And now you end up with 64-bit Word, and you're going to basically compare and swap that somewhere in ah in an atomic hash table very quickly.

Speaker: Okay. Because you need that stuff. That stuff is going to be called by a ton of people very, very often, right? Every time you type check, every time you look up a type, you're going to have 64 cores that' going that are going to, well, 56 more. um They're going going to bang at full speed on this data structure. So it really has to be very fast. so now Yeah. And they're going to be serious hotspots in that. Yes, exactly. Okay.

Speaker: um yeah so you need And so also what you can do is, well, I'm not going to get into too many details. So now you have a data structure that can keep track of all the dependencies very efficiently.

Speaker: And now the aim of the game is to re-walk all the older phases and do all the work that needs to be done incrementally. Right? So, right and the parsing phase is very is relatively easy because you just have to figure out if the file has changed and if the file has changed, you need to reparse it.

Speaker: Great. That's nice and easy. But then in the yeah phase where where you declare the environment, that starts to be trickier, right? Like now when there is a class that changed, you need to walk. So you need to do two things. You must you must first clear up the old environment. So there was an old class called a and you now need to clean it up because this class doesn't exist anymore.

Speaker: And you need to put in your new dependencies. right So you need to put the new version of a And where it's tricky is that you need to to Be smart about the work that you'll have to redo again. right so Figure out what depends on A and redo the right work to basically end up in a place where you would have been had you initialized the code in the state that you're in. Take me through that concretely because I can imagine a class like some logging function.

Speaker: I need to update the AST for that, but also that's going to hit a lot of other places in the code base that depend on my logging function, right? ah Yes. i mean I mean, incremental does not mean magic, right? If you do touch something that the whole world is using, if you decide to change the name of the class int,

Speaker: Yeah, that's that's going to take some time. you know yeah so but But the fact that there's a lot of dependencies does not change the nature of the work. Basically, whenever something changes, you have to figure out who depended on the old version and the new and do the right thing. So what what the right thing means is clean clean up the old stuff and then go and reintroduce the new stuff. And so...

Speaker: That, because it's also happening concurrently, was, ah excuse my French, was a shit show to stabilize. So this is probably I'll allow you one. i'll allow you one curse word on this show. All right, cool. I used i used my ah my my my card then. I use it for that. It's definitely the hardest problem I had to work on.

Speaker: um in terms of with this kind of you know racy racy, non-deterministic bugs all over the place. um it It took several several, many, many months to stabilize. And one key piece of infrastructure infrastructure we had to build to stabilize this thing was what we called them the the dump. So what we do, we added a feature so that whenever...

Speaker: the language server was in a certain state, we could actually dump the whole state of the server into something that was human readable and that was signed for all the different parts. So for example, for each class, I would have all the signatures of everything that I expect, but all in a form that's human readable.

Speaker: So that's basically I'm able to ask the language server, tell me everything you know about the code base right now. And it will dump it in a very predictable form, right? That must be huge. It's huge.

Speaker: It's huge. It takes a lot of time. And so we ended up building that and building an infrastructure that was testing the the incremental updates. And the way that was working is, so back then we were using Git and our Git repo was huge and we would have thousands of commits a day.

Speaker: And Git is perfect if you're going to test something that updates things incrementally because Git works much, much faster than a human being and does all sorts of nasty things much, much faster than a human being. So if you can keep up with Git, you're going to keep up with any human being typing you know on on their files. Right. OK. So the way we did it is we would check out the current version of the repo.

Speaker: And we would initialize the system and then dump the state and then keep that state in a file somewhere. Right. And then go check out um random revisions, usually far away, a few weeks away. So a lot would happen.

Speaker: Right, jumping around the history. Jumping yeah around the history. And every time where we landed somewhere, let the server update and dump the state. Right.

Speaker: yes Okay. So now you end up with a bunch of of commits. For each of those commits, we have a state that's associated with it. And then once in a while, we would go back to a place where we were, dump a second state, and check that they were exactly the same.

Speaker: Right, yeah. And if they've drifted, then something in the incremental algorithm is wrong. Yes, exactly. yeah And that's how we stabilized the system. Of course, yeah. And if you're doing that with Facebook's monorepo, you've got plenty of sample data to work with. Exactly.

Speaker: Exactly. So that was many, many months of sleepless nights. And the frustrating things with with this, I mean, what was good with this approach is it actually worked. But what was frustrating is that reproducing a bug could take hours, sometimes. Days, I don't think I got into days, but hours.

Speaker: And sometimes it was the non-deterministic. In fact, more often than not, it was no not non-deterministic. Oh, yeah. Yeah, because there was a lot of concurrency going on and the way you access the heap is concurrent. And so, um yeah, if you know you have a bug and reproducing it takes you an hour when you can reproduce it and then you try to make a change and of course of course the bug goes away, you cannot reproduce anymore. yeah And that was the kind of that was my life for a long time and and that was incredibly painful.

Speaker: And that's when I was pulling my hair out and was telling myself, there must be a better way. This is not, I refuse to believe that this is the way to write incremental software. There must be a better way than this. And that's when I started my my journey towards Skip, which back then i didn't realize was going to require a new programming language.

Speaker: So many projects, so many people we've had through on this podcast begin with, dear God, there must be a better way. OK, so what happens then? Did you did you start experimenting with the language next? Did you leave Facebook and start working on it? ah No, I was still working on Facebook, but my first plan was to try to make a framework for incremental programming um on top of an existing language. So i looked at different languages. how Hack was actually one of them because it makes sense from...

Speaker: Back then, i looked at JavaScript. I also looked at OCaml. um And i ended up always with the same problem, which was that none of the mainstream languages had a notion of immutability that was strong enough for what you need to do, incremental and incremental framework. So let me walk you through this because it's really not that complicated.

Speaker: right If you're going to build something incremental, there isn't a million ways to go about it. You need to basically maintain a bunch of caches. right You start your program, you cache a bunch of intermediate states, and then the next time around when something has changed, you invalidated only the pieces that have changed and you can reuse all the all these pieces of state that you kept around.

Speaker: And so 99% of the effort is going to be a cache management problem. right So million dollar question. When you take an object out of the cache, are you willing to pay a copy for it?

Speaker: Yes, no. If the answer is yes, then building a framework for incremental computes is going to be trivial and you will be able to do it in pretty much any language very quickly, probably in three to four weeks. It's not going to be that much work.

Speaker: okay If the answer is no, well, now you have a problem. Now you just handed over a pointer to a computation and you need a guarantee that this computation is not going to modify whatever you handed over.

Speaker: yeah And that notion of immutability, unfortunately, is not the one you find in in any language. right So in most languages, when you say const, um you basically say that, so let's say I'm in C++, there's written const. What that means is that this particular function is not going to modify this object. It does not mean that this object can never be mutated by anyone and no one will ever hold a mutable reference to this thing.

Speaker: but So yeah typically, const is going to be a function of the computation. It's going to be a property of the function. It's not a property of the value, right? and so and Yes, someone else might be modifying that value while your constant function doesn't. Exactly.

Speaker: yeah And the other thing is you need the transitive closure of whatever you're touching to be immutable. And that's going to be very tricky with all the parts of the language, all the abstraction that were made to hide um data and abstract away data. So for example, in a functional programming language, your problem is more going to come from closures.

Speaker: which are there to hide values and hide types behind an abstraction. And then it's going to be very complicated when you have a type with a closure. Well, you have to basically you know put your hands in the air like, well, there could be some mutable things in there that I don't control. right and Right. Yeah. In the world of objects. A closure captures an environment that you're not really supposed to be able to get at. Yeah. Exactly. And then in the world of objects, it's going to be, well,

Speaker: the the the fact that you can hide generics, for example, through inheritance. So for example, imagine I have an interface.

Speaker: one, two, three, and an interface called show. So this interface has one one method called show. Well, the problem is anybody could be implementing that, and I don't know what they're hiding in the object. So if I have a piece of code that uses this interface, Well, I don't know what what piece of mutable states could be in there. And the problem is exactly the same as with closures. These are language constructions that were created to abstract away data in function.

Speaker: And they're just doing their job. They're abstracting away the things that that's what they were done for. And you, your problem when you're trying to build an incremental framework is you're going to look at this and you're like,

Speaker: Yeah, but I need this stuff to be immutable. So I need to be able to look into it and have a guarantee that this stuff is immutable. so Yeah, you've got to be able to pierce that veil. I have to ask you, though, because like if we're talking, what the 20 teams, yes, I would agree most mainstream languages didn't have any good idea of immutability.

Speaker: But you've got things like Clojure or Haskell. Did you consider those? I did, yeah. um so The reason why I did not pick up on a purely functional language is because um the APIs were too complicated for my taste. So um let me try to explain the problem. And I have a lot of friends who love Haskell, and they're probably going to send me horrible, hateful messages after they see this. But that's OK. I'll take it. um So ah what's going on is when you... An incremental framework is basically something where you want to keep track of an execution, right?

Speaker: And also you want to keep um the pieces of states that are unrelated with each other as independent as as possible, right? So for example, in our example of Hack,

Speaker: um If you passed a function ah you used ah a function to pass a file on the one hand and another function, well, the same function to pass another file, the states that they produced, you don't want those two pieces of state to depend on each other. If I have a file that changes on one end, i i don't want this to affect the other hand. So you have a lot of states to manage.

Speaker: but you want to keep all those different states independent. The problem is the only way to manage state in a purely functional language is is with a monad. so Either you encode your own monad where You take in the state of the world and you spit it out through the return. If you want to play that game, you can.

Speaker: But basically you are using a monad. You're just encoding it with your little fingers. But there isn't a million ways to, you know, encode states in a world where you cannot change things. The only way you can do it is take states from...

Speaker: the input and then produce a new state. And then you can make that efficient if you can you know the optimizer can figure out that you can do that in place and things like that. So I'm not talking about efficiency here. But the way to do that comfortably in a programming in a purely functional programming language is typically to use a monad and typically to put everything in the same monad.

Speaker: because then things just work beautifully. if In my experience with with monads, not just in Haskell, I've actually used them in OCaml plenty as well.

Speaker: and My experience is that as long as there's just one monad, things are great. It's a bliss, everything works. But the problem is that's actually not what you want. What you want in the context of incremental incremental computing is that you're going to end up with many, many different pieces of state and you want to keep them as independent as possible. So you don't want them to be in the same monad because you the if you put them in the same monad, you're kind of tying them together.

Speaker: Right. Because right now when once this this, so I have my state that comes in this way, i did some work, i it comes back with a new state.

Speaker: If I make this piece of code here depend on that state that came from here, that means that every time this function changes, I'll have to recompute this. And I don't want to do that. So the only way to do that is to introduce multiple monads.

Speaker: And I just found the APIs horrible, untractable. I'm a big fan of Haskell, but I will say for journalistic balance, there are ways to solve that problem in Haskell. I can understand why you wouldn't love the UX of it. You're talking about Monad Transformers? like yeah Yeah, yeah. All that good stuff.

Speaker: ah Yeah, yeah i like I like them, but i wouldn't I wouldn't force them on other people necessarily. Yeah, so that was my take. It was like, I think it's too complex and I don't think it's natural enough. So that's when i was like, I was thinking, well, what's probably going to be a better a better thing to try is to extend an an existing type system such as OCaml.

Speaker: ah to try to add this notion of immutability. And very quickly, I... No, not very quickly. It took me a while to figure out that it was not going to work. um And the reason is, whatever language you choose, if you add this type system on top of this language,

Speaker: on top of the language that you've chosen, you will end up with a programming language that is fully incompatible with the existing language. So let's do it together. So we start with Java, and we decide, both of us, to build reactive Java.

Speaker: Great. So what I need for reactive Java is I need a notion of immutable objects, right? And so I build a new notion of immutable object, I modify the type system, yada, yada, yada.

Speaker: ah So are you talking about like for every string, there's a mutable string for every map, there's a mutable map? Or are you saying there are new primitives in the type system to express that something can be immutable?

Speaker: Either way, it doesn't really matter. So you can choose to extend the existing classes so that they can now become immutable. Or you can create your own thing on the side. If you do create your own thing on the side, you're gonna in either cases, you're going to run into the same problem. So let's say you did that successfully.

Speaker: Well, my question to you is what happens when you take you go from the world of reactive Java and you call normal Java? So I have my immutable object.

Speaker: I want to pass it to normal Java. So either it doesn't work at all. Like one way to go is to say this is not allowed. Okay, well, if this is not allowed, that means that basically reactive Java cannot call normal Java.

Speaker: but at least not without a copy. right Or you let it happen, but if you let it happen, now potentially all your en environments are broken. right Because if you let an immutable... Well, I don't need to to work that out with you. So yeah now we have established that reactive Java cannot call normal Java.

Speaker: And that includes the standard library. so all of a sudden, you have a version of Java without a standard library. that That's not very fun. Yeah, it suddenly looks like you've recreated a whole new language anyway. And then the other way around too. So if you come from Java, can you call immutable Java? Well, no, because the the unsafe world could keep mutable references behind your back. And if you allow that, then potentially you're breaking the system. So the only way to call to and from React to Java to Java is by serializing and paying for a copy.

Speaker: And then if you're going to do that, then you may as well just keep them separate languages. Right. Yeah. I think Clojure came to a similar conclusion in the early days. It's like everything has to be immutable.

Speaker: And if there's an escape patch, the moment you use it, all bets are off. And I disagree with that. I think you can um make a type system where there is mutability, but the the type system has to track the mutability very precisely. And this is basically skip. So it's okay to mix mutable and immutable. And you can have immutable values that are exactly with all the same you know guarantees you would have in Haskell.

Speaker: um that live in a world mixed with mutable values. That's totally fine. But the type system has to track it. And the way the type system works is the following.

Speaker: um Mutability should never be able to hide within a type. so if i have um So, for example, we have two different kinds of closures in Skip. We have flat arrow closures that can close over mutable types, and we have scurly squiggly tilde arrows, and these are provably immutable.

Speaker: um and And so those, they will not let you close over something something mutable. And we have that with pretty much everything. So you have an object. And if this that object can contain something mutable, you will see the keyword mutable before it's type.

Speaker: So whenever I look at the signature of something, I can directly point you at the places that can be mutable in that thing. Right. So is it fair to say, trying to draw a parallel here, you've looked at Haskell and said, here is a type system that's very concerned about tracking side effects and said, well, the ergonomics of that aren't great. Let's just build a system that tracks the specific side effect of mutability.

Speaker: Exactly. And if you do that, then the ergonomics are better, a lot better. Because mutability is not just any side effect. It's it's something that can live side by side. So for example, mutability in many cases,

Speaker: the order does not affect mutability. We just care about the fact that something is mutable, while when you have effects, you care a lot about the order of execution. right um Here, the fact that it's mutability and the fact that the only thing we care about is is this thing modifiable? Yes, no.

Speaker: um That makes the ergonomics... much, much better. And so it actually created a development environment that we tested within Facebook that allowed people to write functional code without realizing they were writing functional code.

Speaker: um Because within the function, they would... So what we would tell them, so if somebody came in with a profile that was not a functional programming profile, we would tell them, look, inside the methods,

Speaker: go nuts, modify whatever you want. But then whenever something has to leave the local scope, try to keep things as immutable as possible.

Speaker: And if something needs to be mutable, you can, it will happen, fine, you can still do it, but avoid it as much as possible. And if you do that, you end up with people who end up writing functional code very, very quickly. they Because what is actually painful is the the whole fold left, fold right stuff when you're used to you know for loops and modifying locals. That is the stuff that painful, right? But using an AVL tree instead of a hash table

Speaker: you know It's a dictionary, right? It's it's fine. so um So that's that was the the spirit of Skip. right where okay we So creating the language was the first step, at least. Okay.

Speaker: Before we move on, give me a flavor of Skip. what is it like What's the syntax look like? How fully featured did it become? Or was it just something for reactive programming? What was Skip like? Skip to program. So you can go on skiplank.org.com. I always forget. skiplank.com. And if you look at docs, you have an overview of features of the language. Okay. I'll put a link in the show notes. Cool.

Speaker: ah But basically the language, um I would say, is... ah It looks like Scala, Java, this kind of language. So it's a language with objects.

Speaker: But um the basic block is immutable objects. ah But what is... ah what is um unusual is that there is no difference in skip between a constructor and an object. They're actually exactly the same thing.

Speaker: So in most languages that introduce pattern matching and are object-oriented, they either have a special class, like a case class, or they have another construction for ats um In skip, not at all. The two are exactly the same. So if I create a base class and I create a bunch of subclasses by extending this base class, I can totally pattern match on that base class directly without having to. And I can pattern match on anything. So okay the the it's really

Speaker: o o meets f p um with um And then the the language feature is pretty complete. We have a an equivalent of type classes, so we have traits, which is how they're called in the OO world.

Speaker: We have... um A relatively simple macro system but that lets you write you know the tedious things like hash functions where you don't want to rewrite those by hand every single time.

Speaker: um So think, derive, blah, blah stuff. We have something like that. um And then, yeah, the language is pretty complete. It's not a toy.

Speaker: I have to pick you up on two things you said there that blew my mind temporarily. The first is immutable objects. I think of objects as almost by definition mutable.

Speaker: They alter their own internal state. How does an immutable object work? It's just that the internal state is immutable. And so you can think of it as basically it's another form of closure. right So first of all, we have mutable objects too.

Speaker: So I don't want to give the the false impression that all the objects, but you are encouraged to use immutable objects. That's going to be the basic block. So let me walk you through how that would actually look like together. So imagine you have...

Speaker: an immutable object and you send a method, ah you you call you send a message and call a method on that object, right? In an immutable world, what would would happen is either the state of the object is not affected, and if that's the case, then you just return whatever you were going return. Yeah, I call user.getName and the state hasn't changed. Exactly. Or the state has changed, and what you would usually do is return a version of yourself with ah with a new updated state.

Speaker: Okay, yeah, so you're just returning a new but ah copy of you. Exactly. If someone tries to mutate you, you don't mutate, you give a new version that would look like mutated version. So we still have also mutable objects that work exactly exactly like in Java, and then you just have to write mutable in front of it. But if the object is immutable, that's how you would work with it. so And then the ergonomics of the language is done in such a way that manipulating those objects is actually very pleasant.

Speaker: So there's not much there there there are a lot of a few operators that makes interacting with these kind of objects very nice. So for example, one that i think is i don't think exists in any other language, so I think that's a true innovation, is the bang operator.

Speaker: So um typically in most programming language that have both um lead bindings and um and um assignments, they have a syntax that's separate for the two, right? They'll have an equal, not always, right? Sometimes it's equal and and whatnot, right? So for example, in OCaml, you have equal and colon equal, right?

Speaker: and Okay. The problem is with us, ah what we wanted was to do two things. um So the way the equal works in skip is that the equal can serve both to be a let binding or modification.

Speaker: If it is a modification, on the left-hand side, you have to put a bang in front of what you want to modify.

Speaker: So right if I go back with my ah with my object, so I just called an object. let Let's say that this object is a map that associates, you know,

Speaker: um ah is it a map? Is that a good? ah No, let's let's let's take an object that returns its own state and another value, right?

Speaker: Okay. So now you're calling this object, it comes back with a tuple, and you need two things. One, whatever the state of the object was, you probably want to put it in the local that was holding this object and modify that local with a new state.

Speaker: Okay, yeah. And at the same time, take the result and bind it to a new variable.

Speaker: yeah So the way you would do that in skip, you would write bang the name of the object, comma the name of the variable you want to introduce equal the call to the method. I'm trying to make this a concrete example. So let's say I've got an object that represents a string and I call append and I expect to mutate the object and get back the new length.

Speaker: a Why would it? No, string is a but is a bad example because you wouldn't mutate the state, right? um What would be a good example where you have both? um No, string, you would just return the the appended file. Okay. trying to um Wait, get set, so that's not going to work.

Speaker: um Maybe a byte buffer if I'm like serializing things to a byte buffer to send over the network. You would probably use something mutable for that.

Speaker: But yes, it's a good example. I mean,

Speaker: can't believe we can't come up with an example of stateful. that's That's how functional we are in our mind. but So let's say you have a dictionary. okay What can you do with a dictionary that both modifies its state and gives you back? Oh, inserts and tells you if the value was inside or not.

Speaker: but Okay. okay yeah yeah yeah Perfect. You have a set, and you want to insert something in that set, but you also want to know if that insertion created a new entry point or if it was already there.

Speaker: So the way you would write that in skip is by writing bang the name of the set. So let's say bang s, comma um has been inserted.

Speaker: equal s dot insert and tell me if it's a, so that's how you would write it Right. And if you left off the bang, you get a compiler error saying this is an immune yeah exactly a mutation problem. yeah Yeah, exactly. And so your equal sign can mix let and let and assignments at the same time. And that is very important for the ah ergonomics of the language because it's going to make these kind of immutable objects much more pleasant to use.

Speaker: And the bang is actually much more sophisticated than that because the bang is also a lens. So whenever there is a path, the band can be applied anywhere in the path.

Speaker: So let's say i have an object that contains another object that contains another object, and I want to modify one of the fields in this one.

Speaker: the way you would Yeah, I always like the example user.address.postcode or zip code. Yeah, exactly. So here you would write, if you write bang user.address.zipcode equal, what you're telling the compiler is you want to modify the local user.

Speaker: And if that's the case, it will have to create a copy for the path um zip code data right right yeah mean But you could go one layer one layer deep. Let's say user was actually a mutable object.

Speaker: And if it was a mutable object, you could write user.bang ah address.zipcode. And if that's the case, it will modify the field within user. so the the The bang, if it's placed totally to the... It works a bit like a lens in Haskell, right? Where I give you something that can modify what's what's at this position, and it will create the right copies.

Speaker: And that makes the ergonomics of of um manipulating immutable objects much, much, much better.

Speaker: because Especially when they're nested, presumably. Exactly. So whenever you're manipulating immutable objects, you are constantly opening up objects and then re-closing them with a new version. like You're constantly doing things like that. I've seen that. And so the ergonomics of that with that with these little operators makes the whole thing much, much more pleasant.

Speaker: Right. So the other thing you said that kind of blew my mind, which you've begun to answer, but I'm going to ask it. um And you've risked the ire of Haskell people, so I'm going to risk the ire of Scala people. here yeah I always thought, shoot me if I'm wrong, I always thought that Scala asked the question, can OO and FP coexist happily together?

Speaker: And I thought the answer from Scala was they can coexist, but not happily. Do you think you found a path to making them coexist happily and nicely to work with?

Speaker: I think so. But I think the challenge that um they had on Scala was a little bit different because they also wanted to be compatible with Java. And so I had the luxury of starting really from scratch. um Then there are other things which was they were motivated by doing research and publishing things and getting, you know, students to, PhD students to publish things that then could be included in Scala, um which is also things we've seen in OCaml, for example. The object actually was the PhD thesis of Vuillons.

Speaker: and things, things of the sort. Jack Garrigue also added things with his students. um And that was not at all my my intention. So I always wanted to keep a language that was relatively simple.

Speaker: I don't want too many complicated features, but not to the point where it's a toy language. So the key things that I just didn't want to compromise on was I wanted... um some kind of namespacing. doesn't have to be a full-fledged module system, but I need to be able to split things and regroup you know files into different yeah and kind of things.

Speaker: I wanted exceptions, and I wanted them to work well. So something like Go was not going to fly for me. or the fact that Rust also doesn't have exceptions. I understand why they don't support them, but for me, that that was a um um and a no-go. I really wanted exceptions.

Speaker: I wanted pattern matching probably the most important feature for me from you know the functional programming world. And also either type classes or something equivalent. So I tend to favor type classes over functors, even though I understand that functors are strictly more expressive.

Speaker: I just find them less pleasant to use. um So when I write code, because the problem is if you don't have either functors or type classes, like you you and type classes i include traits and what comes from the old world, right? It's all the same stuff.

Speaker: um If you don't have either of those things, then writing generic algorithms is going to become really... painful, right? You will end up rewriting the same stuff over and over again, right? Because you just don't have the the right abstractions.

Speaker: So right I wanted a version of that that works. And I picked traits, meaning type classes. So my OCaml friends are going to hate me now, just because I used functors for a while and I just don't find them as pleasant to use as type classes.

Speaker: Can you break that down for me? Because I think some people listening won't know what a functor is. I come at it from very much a Haskell background, and I think the OCaml world has their own sense of what a functor means in practice.

Speaker: Yeah, so a functor, basically, you take ah a module, so you can think of a module as a bag of types and functions, right? So right think of it as a unit defining types and compute, right? Function. It can be records, can be anything you like. And what you're going to do is you're going to have a meta function that, given that module, is going to produce an implementation of something. So concretely, imagine I want to build a hash table.

Speaker: i could build it with a functor where I ask the user to provide me the type of a key. the type of the value, a function that will know how to do the hash, so given a key is capable are producing an integer, and maybe a few other things, like maybe some hints on how to grow the size of the hash table or things like that. But basically, I asked the user to provide a few types and a few implementations. And given those few types and few implementations, I will be able to produce

Speaker: um um a functioning hash table. And the way it works so the way it works is it's really a meta function. That's what it is. So you by applying a module to a functor, you end up with a new module, and that module module is going to be the implementation of your hash table, right?

Speaker: Yeah. So you're making it concrete from generic hash table from K to V to a specific hash table of int to string. Yeah, exactly. yeah yeah And so my problem with functors is that in this particular example, it works well. But what ends up happening is I end up having the same problems I have with functors as I have with ah deep class hierarchies, which is that um when it's nice and simple like that, when you know you just have one layer of indirection, it's really nice and easy.

Speaker: right But what's going to end up happening is, let's say you have your generic algorithm, but your generic algorithm is not you know something that's going to fit in a standard library. It's actually much more complex and it's more like specific business logic that will have different pieces and whatnot.

Speaker: And so now what will happen is, you will end up adding functionalities and you will need probably more than one module and you'll need probably more than one function to build these things. And now you're adding layers of indirection, layers and layers of indirection to figure out. And in the end, once you have a lot of those functions all over the place, I find the code difficult to follow.

Speaker: So how it so I don't want to say... i mean what I'm not saying you cannot have a good code base with functors. That's not true. You can totally have a good code base with functor. But i would I would put it in the same category as inheritance.

Speaker: Inheritance is a powerful tool. If you abuse it, you really can and end up with a code base that is very difficult to navigate because you constantly basically are trying to figure out where's the meat, where's the actual function doing things, right? And you find yourself constantly following different classes that just call ah more abstractions and even more abstractions. And you you have a hard time you know following what the code is actually doing.

Speaker: I have a similar feeling with a code base that has too many factors because then I look at and I'm like, okay, where is that implemented? And then I have to follow which module it comes from and like, okay, but who is actually, in you know, instantiating that module? And I have less this feeling with type classes because with type classes,

Speaker: it's both the strength and the weakness of type classes is that it's all directed by a type, right? So with a type class, I'm going to say, i have a type. I don't know what this type is.

Speaker: But if it implements this plus minus multiply, then I can do algebra with it, right? Or things like that. Right, yeah. And the big difference between a functor and a type class is the type class actually the The one driving the the what what functionality we're going to use is the type, right? So if I pass an int or a type that implements the right functionality to, ah you know, um something generic that was, it's going to work. So, and the fact that the type ah so it is kind of the the thing grounding where the abstraction takes place.

Speaker: So it's both,

Speaker: I think in my book, it makes it much easier to follow because whenever I want to know if I can use this type with this kind of thing or figuring out where the the where the type class was implemented, I just have to look up the type. So it's very easy for me.

Speaker: But it's also the limiting thing. So the the the problem with type classes, and and I think that was the... the the thing they attempted with implicit. The problem with type classes is that it forces you to have one implementation for one kind of um one kind of ah type class. So you have a type class. You have defined plus.

Speaker: If you have defined plus and integers, you won't be able to have different plus different implementations of plus depending on the kind of abstraction that you want to use.

Speaker: and Does this make sense? I see you frowning. I think you just lost me there. Okay. so think that again So let's say you have a type class and this type class asks you to implement the method plus.

Speaker: yeah Let's say you have another type class that also asks you to implement the method plus. like um Except that those two pluses should do different things on the different type classes.

Speaker: That's going to be a problem in Haskell if you want the same type, but with different implementations of plus to implement both. you will have to create a new type, basically.

Speaker: And Implicits tried to solve this. So that was the whole idea of Scala. The idea was we were going to let you define multiple methods, multiple implementations for the same type.

Speaker: And then based on the typing environment, we're going to choose the right implementation. ah That was the idea of Implicits. Right. Right. so And functors don't have this problem because functors, the type is not what is directing what's going to be um what's going to be instantiated. It's you who do it explicitly with functor call. So you explicitly say, I want to use this abstraction with this module. And in that module, you can write choose whatever plus implementation you want. right

Speaker: Yes, yes. I think I'm with you there. right. So I wanted something that was but basically i as expressive as either functor or type classes. And I settled with type classes because I think they're just...

Speaker: much easier to use, even though they have this weakness. This one weakness that is that if you need multiple implementations, different implementations with with the same type, it's not going to work. You'll have to create different types. And I didn't want to add implicits because I don't think they're a good idea, but I don't know if ah that's enough geeking on types for now.

Speaker: Right. I'm sure we can find some Scala people that think implicits are a bad idea too, but perhaps let's move on. yeah so So you were writing or co-writing Skip the Language at Facebook.

Speaker: That got open sourced. Yeah. And then at some point, you then go into reactive frameworks on top of that. Yeah, so the problem is that um skip the language by itself is only a language that gives you immutability guarantees.

Speaker: But it's really not something that you where you will be able to build a reactive system with it. There was many, many pieces that were missing. So the fact that you have a programming language that gives you the immutability guarantees that you need,

Speaker: That's great, but that is just... the first brick in a much larger ecosystem. So if we go back to what we were talking about earlier, where remember this system where we had this hack pipeline. And if you want to build something like that, you're going to need, you know, things that keep track of your dependency, things that can store objects, things that can um garbage collect memory that can. And so that's what we set ourselves up to do when I left Facebook. Um, So first I spent one year doing research on AI um at FAIR and then and then I left the company to... um I wanted to to know what what the AI was. was fun so this one Which year was this? Because AI has changed a lot in 2018. Okay.

Speaker: And then i I went and built SkipLabs. Okay. okay So you having it's been outsourced from Meta.

Speaker: So you jump ship and you start developing the language or start developing the framework on top of it? No, the language was more or less ready when we left. We have done a few things, but mostly maintenance. So the compiler is stable.

Speaker: It was used at Facebook for several years and it was already stable there. There will be so very, very small bugs, but... And the language itself, I i don't want it to and evolve too much. I really would like to keep it a simple language.

Speaker: um So what we have now is probably what going to keep in the foreseeable future. And so then we started ah focusing on the infrastructure. And so...

Speaker: The infrastructure, so we ended up first building a file system. So remember that heap that I told you about earlier? Yeah. So the first thing we needed to build was an equivalent of that.

Speaker: And um we didn't want everything to end up working like a language server. And the reason is language servers actually take up a lot of memory, right?

Speaker: And there are many tools that where you would want them to be incremental, um but you don't necessarily want to pay the RAM um all the time. you know you you want So let's let's imagine you are working on 20 different Git repos, right? You don't necessarily want to have 20 language servers running all the time, right?

Speaker: And we figured we we didn't actually need that. All we needed was... we wanted something that worked more like SQLite. So you define a file, and in that file, you are going to put all the artifacts of the of the long-lived objects that you want.

Speaker: right And then whenever you call the tool, what the tool is going to do exactly like SQLite, it's going to mmap this file in memory, and it's going to work with with those objects and and do the work. and so now In a language server, the way it works is you send a message to the language server, it does some work, and it comes back to you with an answer.

Speaker: In this world, what you do is you call your tool like a command line tool, and you specify a file where you're going to store the artifacts. And then what it does is it runs your commands, it updates the artifact the artifacts, and and then the command finishes. So it's a much nicer workflow for people who want to use tools.

Speaker: Right. Does this give you a solution to the whole persistent state problem? there you don't have When you restart the system, you have to repopulate the entire cache? Yeah, so that's the thing. So the the cache is actually living in the file. And yes, that solves the persistence problem. But it comes with a bag of problems. So one big problem is concurrency. So you want to be able to have multiple people using the same file at the same time.

Speaker: And so now you have ah read readers, writers, concurrency issues. um And you also have a memory management problem because now you have a heap to where you need to manage not only the objects that live in memory, but you also need to manage those objects that will live in that long-lived file.

Speaker: And they have to form like coherent set. um set Right. So that's the first thing that we worked on. And um to make it scale, so we we ended up with something working relatively quickly. So we had um this file system that was basically an object system or persistent object system, whatever you want to call it, okay that was working. And we had a framework that allowed us to... um

Speaker: ah build pipelines of executions like the one I mentioned when I was describing how hack was working, except that in that pipeline of execution, everything is safe and you don't have to manage dependencies by hand. So if I go back to my hack framework, to my hack pipeline, the way I would write it in skip is I would take my files,

Speaker: And then I would say map the parsing, and I would write a parser for it. That's actually how the the skip compiler works because the skip skip is written in skip.

Speaker: um OK. And so so first you would parse the files. Then those files, without you having to say anything about it, they're going to end up in the artifacts in the in the file system.

Speaker: then ah for them type declaration, you would probably have um a mix of two things. so we have two main constructions. one One that is a map that lets you map over a collection of objects.

Speaker: And another one that we call... um a lazy map. Well, no, we call it a lazy collection, sorry. And that is basically, um that is more like a traditional cache. So you ask for some data, it will do some compute, and when it comes back, it will store this data in the cache. And the two actually work in tandem, and they're they're usually really useful together. But basically, long story short, it's very easy to take a reactive programming pipeline.

Speaker: And by using the skip framework, you write it very naturally. And what I've done back then at Facebook, I could do without all the headaches. I can just make it work and it just scales. You built the whole infrastructure underneath to yeah to make that kind of problem easy. Yeah. Yeah. Okay.

Speaker: um Tell me just give me a bit more detail on lazy collections. So what is that kind of like ah a compute on demand cache? Yeah. So the problem that you have is that there's two schools of thoughts when it comes to um reactive programming. So first of all, there's a lot of debate around what reactive programming is.

Speaker: And some people use this term for streaming because they think that using streams is the right way to model an incremental program. I don't think it's the case, but i don't okay I'm not going to you know ah blame anybody who does. For me, the stream is not the right abstraction because the stream um ex exposes just too many problems. um basically The problem with a stream is that the moment you have one stream, you have a ton of concurrency problems. You have a ton of, well, especially the moment you have two streams. So you have two streams.

Speaker: What if one is faster than the other? Okay, well, maybe you slow that that one down. Well, what if you can't? Well, you add a buffer. Well, what if that buffer fills up? And Those kind of questions, basically, you will have for every single time you um you manipulate a stream, you will have to ask yourself all those questions. And it ends up being a nightmare of of concurrency, basically. So that's that's my take on you know streaming for incremental compute. but Does it make a difference if you're doing push versus pull-based streaming in that example? It does, and most systems let you do both. so For example, if you come from the ah Rx school, the pull version, they will call it ah called a cold stream, and the push version is a hot stream. um and so

Speaker: Yeah, it's and it's actually funny that you mentioned that because um it's actually the true score of thought that i wanted to talk about. So there are people who believe that um the the incremental system should be pull-based. So, for example, if you look at something like Samba, the language server for Rust,

Speaker: This is definitely a pool-based solution where um a user issues a request and basically you're going to do work to try to figure out, you know, to answer to this request. And when When you're done, you ended up using a bunch of caches to you know to respond to this request. And those caches is is going to be the basis for for the incrementalization of what you're doing.

Speaker: right And then you have the push-based people. And the push-based people tend to be streaming people. But we ended up, the the problem with streaming, i think I just mentioned, is is come with ah with a bunch of concurrency problems that you don't necessarily want to deal with. And so we ended up with um two solutions. So one that's push, one that's pull, but they both use a very similar abstraction.

Speaker: And so they're both collections. One is a push collection and the other one is a pull collection. Okay. But the the difference between the two is the push collection is going to be maintained eagerly up to date at all times.

Speaker: And the pull collection is um maintained lazily. So when should you use a pull collection? Basically, whenever you can.

Speaker: It's always better to use a pull collection because pull collections can actually decay in memory. They cost a lot less in memory. Because in a pool collection, if something is not has not been used for a while,

Speaker: What you can end up doing is just dropping it from the cache. and then if so Because you can always recompute it. Exactly. Now, with push-collect, the problem with the only pull solutions is that you're going to have a problem every time you need an index, every time you need to do a lookup. So, in fact, let's go back to my example in Hack where we built a typing environment, right?

Speaker: So let's say we built this this um this environment lazily. like So I don't have a phase ahead of time that's actually resolving all the types and knows what the type of each function is. So I'm writing my algorithm for my type checking.

Speaker: I hit a function. Well, I don't know where this function is defined. but So if I don't know where this function is defined, well, what can I do? um well, I have to go and pass a bunch of files and figure out where this function is. But now if I do that, I can do that, but i have created a dependency from one function to basically the whole code base.

Speaker: i Right, yeah. is Is that unique to the situation you faced with ah dynamic loading in PHP? Because in a lot of languages, the import statement will tell you exactly where to look.

Speaker: ah Uh...

Speaker: So the import statement would typically still not give you a file. It would probably give you a package name. it It depends on the language. So some languages are going to give you a file and then it's nice and easy. But there are many languages where you're not going to get a file. You need something that helps you resolve what's going on. But okay that point is is more anecdotal.

Speaker: What's more important is... Basically, the only place where you cannot be poll-based is when you need a dictionary, when you need a lookup table.

Speaker: A lookup table is typically an index, is typically something you compute ahead of time. If you don't have your index and you try to compute it lazily, whenever somebody is going to try to access the data, that just doesn't work.

Speaker: so right So, yeah, go ahead. No, I'm just wondering if we make this concrete, like if I were building a language server for a language, there would be times when the cursor in my editor sends, what's the documentation for this function or something.

Speaker: There'll be also times when the file system says these files have been written to you might want to update them. So when do I use push and pull? What's the flow there? So you don't worry about push and pull.

Speaker: You worry about collections and you worry about making them lazy or not. And the difference between in in your world, in the skip world, you start with an input collection, which is typically going to be files. They could be other things. Once you have an input collection, you ask yourself two questions. You can either map over this input collection to create a new collection.

Speaker: And if you do that, That is going to be eagerly maintained. So every time a file changes, your your stuff will maintain all automatically for you.

Speaker: ah Or you decide to create a lazy collection. And if you create a lazy collection, it will end up like a cache. And every time you compute, we will... keep around the the intermediate values, and all the cache is going to be managed for you. So if the values are not used anymore, they'll be evicted, et cetera, et cetera. And that's it. That's the only the only two things that you worry about. um Then we derive a reactive system from that code.

Speaker: So it works a little bit like React.js, where you write the code pretending that time is frozen. You never think about events. You never think about updates. For you, you you in when you write your code, you imagine that you have these inputs and your job is to write the initialization.

Speaker: That's all we're asking you to do, right? So you pretend you're initializing the system and that nothing will ever change. What do you expect to see, right? And... um Once you've done that, if you've used the primitives that we gave you, we can derive a reactive system from it. So whenever one of those files changes, we will know what to recompute exactly. We will be able to you know update efficiently the system.

Speaker: Right, right. I think it's probably a good idea if we talk about how this gets used, because obviously there's an example of like incremental computation of large code bases, which not everyone will face. I can give you a personal example. I was playing around with Skip, and um I just wanted to kick the tires on it so I know what we're going to talk about. right And I thought, well, I've got Claude running in several directories and Claude keeps an append-only log of all the messages it sent.

Speaker: So I built this thing in Skip, which watches all my Claude project files and reacts to them changing and displays what the latest message in each session is on a web page.

Speaker: Cool. Very cool Which is kind of fun. Yeah, I i enjoyed that. But like where else does this get used? like What's your user base using Skip for? So, I mean, what you're talking about here is the version of Skip that has been ported to JavaScript, right? Oh, okay. I'm jumping ahead. Okay. Yeah. I suppose, right, that that's the one you used.

Speaker: yeah Yeah, yeah, yeah. Okay. So if that's the case, then for this kind of project, the people who are going to be interested in that is people who want to build real-time systems, right?

Speaker: um And ah basically, collaborative stuff, real-time stuff. You have a bunch of things that need to update whenever you know you receive a log. I mean, what you what you wrote is a perfect example of the kind of things that you would write for the skip framework.

Speaker: However, if you're using skip the language, then the go-to problems are typically going to be tooling. So imagine, let's say you want to build something that generates I don't know, generates docs, right? So you have new version of Doccision that you want to write and you have all sorts of ideas on how to do that. The problem that you're facing is that every time you want to update the talk the docs, it takes 10 minutes and everybody hates you. And yeah, now if you write it in skip, you can make it such that it's both going to be fast because it's all going to be parallel and fast.

Speaker: and super fast. And also, it's going to be incremental. And because we have this file-based state thing, the incrementality can be made across machines very easily. Because you can totally have a circle CI with a base state for your um for your for your tools so that when CircleCI starts, it does not start from zero. It actually starts from the state that you saved for your CircleCI box. And so right now you can you can get huge speedups.

Speaker: Yeah, I can totally see that. Regenerating the docks every time when you notice a small change. Definite pain. Yeah. Okay. so So I have jumped ahead slightly. if If someone were building that, would they build it in skip?

Speaker: Or should we start talking about the the port? I think you can. you cant i would build it in Skip. You can build it in JavaScript too. But in JavaScript, it's more going to be like a language like a language server. like it's It's not designed for the tooling use case. So I would say for ah for tooling, command line tooling, I would encourage people to just use Skip.

Speaker: Okay, so then let me ask the question that bridges to that. Having built your own language, why then say we're going to port this to good old TypeScript? ah It's just to make it more accessible. So what was happening is um i think we there's two parts in a reactive system. There's the logic itself.

Speaker: And then there is the framework. And we have a very efficient framework. And it's true that if the logic itself is not...

Speaker: is ah is suboptimal because JavaScript is not the fastest language out there, but it's definitely not the slowest out either. We thought that we would be able to expose what we do to much, much larger crowds without you know having to educate them on a programming language. So basically we did that because we wanted to lower the bar to use our tools.

Speaker: Okay, that makes total sense. Does it end up making trade-offs because it's not in your in your mutable, immutable language? Yeah, so I mean, and yes, the the biggest trade-off is around immutability. And so the way we did it is...

Speaker: ah we need all the objects to be in our heap. um And um the the the big question was, how do we expose those objects to JavaScript? And so what we ended up doing is use a proxy object. So JavaScript has the notion of a proxy object. So it's basically an object that to the user looks like an object with fields. But in fact, when you access those fields, you're accessing a getter and setter of your choice, of your implementation.

Speaker: And so what we ended up doing is building proxy objects that would go and look at data in our heap. And so this way, right we are exposing what looks like JavaScript objects to a JavaScript developer, but in fact are not, and all the underlying data lives in our world.

Speaker: Right. So is it then doing when necessary, it's doing the copy of the data so that it can be mutable? Yeah, exactly. So there's a graph of computation. And so what happens is, um let's say we want you to parse something. right So we're going to give you the data to parse. And for you, this data is going to be a proxy in data that lives in the skip heap. You're going to do the parsing. You're going to finish with an AST.

Speaker: And what we're going to do with this AST is we're going to copy it back into the skip heap. And then we're going to wipe out all all the memory. Because what we don't want is to have a ton of memory that lives in the JavaScript runtime.

Speaker: Because if you do that, the performance of the GC is going to degrade. So you want all the data to lives to live in in our our heap. okay okay

Speaker: I'm wondering where we should take this next. I want to just check what's the TypeScript framework written in? Is it written in TypeScript? or Are you compiling skip down to WASM or what doing?

Speaker: ah So we have two two ways. So Skip actually compiles also to WASM. Okay. And so we have a version of it that works in a browser, but we also have a native version that is more efficient.

Speaker: And so if you're using TypeScript in a browser, you'll have to use the WASM version, and then you'll have a limit on the size of the heap that's going to be basically comes from the browser. It's a limitation of the browser. You're limited to one or two gigs, depends on them depends on the browser.

Speaker: And then if you're running server-side, then we have native bindings for Node and Bun. And okay in that scenario, you're actually using... So Skip compiles to Wasm. It also compiles to native. It also has a very interesting memory model, which I think is pretty unique.

Speaker: um and allows you to have um predictable latency for the changes. Because one of the big problems when you when you write something real-time or something incremental in a garbage collected language, one of the big problems you're going to face is the garbage collector pause times. So if you want to build something real-time, if the GC kicks in, it's not so much real-time, right? Yeah, your latency just goes into a magic hat, right? Yeah, and so we don't have this problem in Skip because we were able to use the guarantees of the language to... In fact, there's two things that we were able to do very differently. One is the garbage collector, and two, it's how we manage concurrency.

Speaker: um So first, the garbage collector, the typical... um So it depends who you talk to, but the typical person... who comes from an imperative background, they tend to underestimate the problem. They're going to tell you that reference counting has basically solved the problem or that the Rust memory model is the right memory model for everything.

Speaker: And okay that's because they are not exposed to, or they don't write very often, very functional code. But if you write very functional code and in a reactive setting, you have to write very functional code because you're manipulating a ton of immutable objects. And so your style better be immutable because otherwise you you're going to have a ah hard time.

Speaker: So if that's the case, then the the patterns of allocations are very different. You're really much, much better off in terms of performance with a garbage collector, especially a generational garbage collector.

Speaker: Because what you'll do in those kind of setups is that you will end up allocating a ton of tiny objects with a very, very short amount. ah lifespan And so you want to collect them quickly. right And so okay the the right approach in terms of performance is to have a tracing collector.

Speaker: Now, the problem with a tracing collector is that it's ah a generational tracing collector. is that It's basically a ticking bomb, right? Like you are accumulating you know garbage that you push push into the old generation, but eventually you'll have to to garbage collect the oldest generation, and that's when you know all bets are off. It could could take five minutes, right?

Speaker: Right, yeah, yeah. So we are not in this situation in Skip because... Our heap is a heap of immutable objects um And the so there's the heap of persistent objects, right? So this is where the lion's share of the data is going to be. And those objects are provably without a cycle. That is also a property of the language. So now we have immutable objects that we know don't contain cycles.

Speaker: So those objects that are long-lived, we can actually reference count them safely, right? Okay, yeah. Then when an update comes in,

Speaker: What will happen is I'm going to simplify. It's not what's actually happening, but so that people get it for the VM. You can imagine imagine memory is just leaking.

Speaker: So let's say you are in a language server. Somebody changed the file. okay So you take the new file, you start parsing it. And let's say you're parsing everything just leaks memory, just a bump allocating and just keeps on going. right okay So now eventually it's going to finish.

Speaker: It's going to finish. And when it's finished, it's going to produce a new heap because the heap is immutable. It's going to produce a new version of the heap with that contains objects that were just allocated during this phase.

Speaker: Right? Yeah. All right, so now that's when you enter commit time. So at commit time, so I'm going to simplify for now. We're not going to talk about concurrency. So we're going to imagine that there's only one person working with a language server.

Speaker: So if that's the case, then at commit time, you have an old heap, you have a new heap, And all you have to do is reconcile the two. So basically, copy the stuff that survived in your bump allocated data structure and malloc it. It's not a malloc, but you get the idea where the reference counts and free the objects that were that are no longer necessary in the old heap.

Speaker: And the freeing of the objects, you don't have to do all at once. So if your system is busy, you could delay the moment where you free them. Now, what's interesting with this algorithm is that the complexity of the GC, so of course it's not how it works exactly, but you you'll see where i want to where where go with this. The complexity of your garbage collector garbage collector is big o of what you have allocated.

Speaker: It's not big O of the size of the heap. like i have allocated I had a file that changed. I allocated a bunch of objects. The worst case scenario for me is all those objects survived and I will have to copy them into the main heap.

Speaker: Right. Right. So you're just garbage collecting the abandoned changes. Have I got that?

Speaker: No, the key idea no is because the persistent objects are immutable and reference counted and are provably immutable and provably reference counted, when a change occurs, I can operate under another system where I'm basically leaking memory.

Speaker: And then once my update is finished and I want to reconcile what was just allocated with what was living in the long-lived heap, the worst the the most time it can take to do that is if all the objects I just allocated during my update actually survive.

Speaker: Okay. And so the reason why this changes everything is whenever you're building something real-time, something where you need you know you needed to respond every second, every 10 milliseconds, whatnot. Well, now you can measure things. You can measure if your update was short enough.

Speaker: And if it was short enough, it's always going to behave this way. It's never going to kick in into this major collection mode where it has to now garbage collect the whole heap that is potentially gigs and gigs of data.

Speaker: I want to make sure I'm getting this right. so And I'm not 100% sure I am. um So I allocate a load of stuff. Perhaps I leak memory. And then I want to copy it across into this immutable reference counted heap.

Speaker: So i'm just I'm just copying those things into a reference counted world and then generationally garbage collecting the stuff I leaked. The stuff you leaked, you throw it away at the end of the yeah update. And of course, you don't have to leak. You can use a generational collection for that scheme. But the generational collection is going to be very efficient because every time you hit something that is in the main heap, you know it's provably immutable. And so you can just stop tracing it. Yeah.

Speaker: Okay, so is it really there's a kind of the two strategies for managing memory? There's the long-term heap, which is reference count. It doesn't need a garbage collector. Exactly. And then there's the ephemeral stuff, which does have a garbage collector, but it's generational. It's mostly just throw it away. Exactly.

Speaker: That's exactly Yeah. Okay. Okay. Yeah. Yeah. I'm with you. And so now because of this distinction, you never have those long pause pauses because the the short-term garbage collector only operates on the data that was allocated during the update. And so that that is never gets very big.

Speaker: Right, yeah. The amount of work is based on the amount of new data you care about, not on the amount of computation it took to get there. Yes. And that's predicted to be small in most cases, most of the time. Yeah, I mean, if you want your update to be fast, it will have to be small.

Speaker: Because if you want something big, then it will take a long time to compute it, and then you're probably not going to care all that much about latency anymore. Yeah, but what you'll be able to say then is that the garbage collector pauses would be proportional to the amount of data you explicitly as a programmer say, I want to store that.

Speaker: Exactly. and the worst case scenario Not proportional to some unpredictable framework thing. Exactly. it's going to be The worst case scenario is going to be 2x. Basically, everything you knew you allocated is going to survive and your code was...

Speaker: 100% dominated by allocation. And then the worst case scenario is 2x, basically. Okay. so that That makes sense. And then for concurrency, so i remember we had ah an old heap. and it comes in a new heap. And now we have to imagine that there's more than one user. And so what's nice and easy is when you're trying to make a commit and the heap has not changed.

Speaker: So if that's the case, then you just flip the new head of the heap and you're done. Now, imagine the new the the heap that you're trying to commit to has changed you know in the meantime.

Speaker: So what typically you would do in you know those kind of multi ah multi-revision systems is that you would have some kind of rollback. But we don't need a rollback because we can reconcile changes after the fact.

Speaker: and So that's a really interesting way of dealing with concurrency. So okay imagine I started with a heap and then i do some work. But then when I come back, the heap has changed.

Speaker: Well, it turns out we have a way to know very, very efficiently which parts of the heap have changed. And so that is super fast. Log n. We will get this this information very, very quickly.

Speaker: Now we get those changes, so we know exactly which part of the heap has changed between now and the moment in the past when you started your transaction.

Speaker: Right? Okay. So since we know that, what we can do is tell you to update incrementally. After all, you're an incremental system. You're trying to you're trying to bring in you know a change, and that change was computed through an incremental computation system. Well, all you have to do is to do that a second time, but this time under a lock, because you don't want... the heap to change again, because otherwise if the heap keeps on changing under your feeds, you're going to have a fairness problem. So here's what you do.

Speaker: You take the lock, you say, okay, what has changed between the moment when I started my transaction and now? You take those changes, you feed it into the machine by saying, now go update incrementally.

Speaker: and will incrementally update the work that you've done, which is not going to be from scratch most of the time because you will be able to keep a lot of the stuff that you computed. Now you end up in a new state and now you commit that.

Speaker: Okay. Yeah. Yeah. I wonder if you've ever considered just making that single core. Would that not... Because that would solve the locking problem.

Speaker: but would Would that just throw away too many other advantages? ah Yeah, no, I don't want it to be single core. There are just so many cases where you want um you want multiple readers, multiple writers. i mean, so many systems where this is really what what you want. So I i mean, i considered it ah for a while. And then so one of the things is that we built a SQL database on top of... um in Skip. And the the reason why we built that SQL database was we wanted to test um the quality of our file system and all the persistent stuff that we've done and compare state-of-the-art databases and see how we were performing.

Speaker: It turns out we're performing very well, so very happy about that. It's called SKDB, but we have we We retired it as a product, so we're not trying to sell this as a product anymore. We thought for a while we could make a product out of it, but turns out not to be a very good idea. But yeah, we have a reactive database, which means that it's a database that lets you... um ah write queries that that you can then listen to. so You have a SQL query, you write a select, so and then it will come back with the data.

Speaker: and Then the the query actually stays live and tells you what changes are coming in the database that affect your query. and It does that incrementally. Okay. so You can get something like a subscribable view. It's exactly right.

Speaker: It's subscribable queries. And it's actually a clone of SQLite. So we took all the test suites of SQLite and used that to to make sure that our stuff was working. And it's it's at around that time that we spent a bunch of time on getting the concurrency right, getting all the...

Speaker: multiple, you know, we like TPCH and these kind of ah these kind of um benchmarks are designed to stress stress test your your database.

Speaker: Yeah, I can believe some stress was found during that period. Yes. Okay, so out of curiosity then, why did that not turn out to be a good product, even if it was working?

Speaker: So it's not a good product because um the rest of the logic... So here's the thing.

Speaker: Let's say you do, so first of all, people don't want to switch databases easily because it's a big decision and they don't trust you. yeah So that makes perfect sense to me. And if they're going to switch databases, there needs to be a huge win for them.

Speaker: It has to be you know a life change a game changer for whatever application they're writing. In our case, so the problem was the following. You come with a reactive database,

Speaker: And then they tell you, yeah, that's cool, but the rest of my logic is not reactive. So now I have this database that's spewing out, you know, um ah changes. And then what do i make of those changes? Right?

Speaker: how do i how do i do I just rewrite all my backend to be in streaming? So then what we would tell them is, well, you can write more logic in SQL and then it will naturally fit into the system. And universally, the answer is well, we don't want that. We want to write less SQL, not more.

Speaker: So if to use your thing, we need to write gigantic SQL queries. That's exactly what we don't want. So thank you very much. We're not interested in that. So that was the main concern. And that's when the idea of building...

Speaker: a JavaScript framework made sense, which was, well, people who want to build reactive system, they will want a full-fledged programming language to be able to write those systems, and SQL is too limited.

Speaker: And so that's when we engage with ah with a TypeScript framework. Okay. So let me... ask for a similar result in a different way. because um so My thing I built is basically just using the file system to notice when files change.

Speaker: and It's just rereading the last few lines of each file. But something I've always wanted is I can look at Postgres and get a live set of updates coming off that query and stream them out to a client.

Speaker: yeah So do you have some mechanism for integrating with a more traditional database like Postgres? Yeah, we do. So in fact, we have a postre Postgres plugin. So it's actually really easy to build um um subscribable views on top of Postgres. We have everything in place to do that.

Speaker: So you hook up to Postgres, but it's very transparent. It will hook up to PG Notify. And then it will give you a view of that data in ah in the form of a collection in the skip framework. And then you can do whatever you want with those collections.

Speaker: And every time Postgres changes, the whole chain will update you. So that's doing something like running the query once and then using pgnotify to get table change notifications? So we we don't, on all the right path, we don't touch, right? In the the skip framework. In the skip framework, the only thing we do is we... we um oin help you with a read path.

Speaker: So if you have a read path that you want to make subscribable, so in step one, we'll take the data that you want to manipulate from Postgres to um lu to JavaScript. And then we are also going to add the right triggers. I mean, it's not called a trigger. It's PG notify stuff to get notified whenever the stuff that we looked at changes.

Speaker: And then, um well, there's an initialization phase where we get a first run of the data. And then the whole thing is plugged in automatically for you so that whenever whenever you write something to Postgres using a Postgres client, Postgres is going to make the write.

Speaker: When the write goes through, there will be a notification for us. We get that notification and we update our system. And the notification contains all the data you need, or you go back and rerun the query?

Speaker: The notification, i i don't know. I'm not the one who wrote that. I'm not the one who wrote that. I don't know if it doesn't go back to know what is in there.

Speaker: Yeah, I don't know enough. I can't answer this question. I'll just have to link to the docs for that one. Yes, Okay. um so Let me just track this forward. We're probably overrunning on normal more time, but I'm finding this really interesting. so if you've got the time, I'm carrying on. and so so I could hook up to my database of choice like Postgres.

Speaker: get streaming updates into a reactive system. I write my computations in TypeScript. What happens when I get to the server a boundary with the client? Is that doing like a WebSocket or something?

Speaker: So the server boundary with a client, um right now, that's not something that's... So, yeah, I mean, right now we use...

Speaker: server-side events for that. okay Because the the thing with WebSockets is that they're not, I mean, they're used and they're definitely supported, but they're not supported everywhere.

Speaker: And so the problem is you have a lot of firewalls, you have a lot of you know network-y stuff that can happen between your browser and whatever whatever you're trying to whatever machine you're trying to hit. And we found that what's nice with server-side events, because they're basically just HTTP, they just work anywhere, and there's one one direction that we don't need. So the thing is, the way you establish a subscription is you call a request,

Speaker: And then once this request has responded to you, the subscription always goes this way. The server will push things to you, but you you don't have the right to actually send anything to the server. So we don't need the communication in the other way. so You're just a broadcast system.

Speaker: Yeah. so So for us, server-side events was a much better fit. And because it's unmodified HTTP, it just works everywhere. Okay, yeah, yeah. and is it And it's just sending deltas for each, every time there's a change? Yeah, so it's changing deltas and the...

Speaker: so Actually, no, it's sending deltas internally, but to you, it shows you the entire object. But if you need something more fine-grained, we could help you hook cook it up. Because under the hood, what's actually sent is only the deltas, but then we reconstruct the object for you so that you don't have to do it on your end. But if you want something more specific, we could do it.

Speaker: Are you saying like if my view is 200 objects and one of them changes, I'll get 200 new objects through? or'll So it no, if your view is a collection, then no. ah But if your view is an object, I believe, yes, you get you get the full object every time.

Speaker: Okay, right. So you don't you don't do kind of those in-place lenses we talked about earlier. No, yeah, we don't do that. But you could. Okay, okay, right.

Speaker: Okay, so having tried my example with Claude, I can sort of see the pipeline from some back-end watcher process to the client.

Speaker: But I wonder if this gets used with things like maintaining live context for LLMs. That might be a popular back-end use of this these days. Yeah, i mean the stuff i mean, for LLMs, we haven't seen much use. The stuff that is popular with popular.

Speaker: for a definition of popular, yeah collaborative stuff and um people who really care about latency client side. So imagine you're building an experience where you want your your app client side to feel really snappy. Like you really don't want to pay a network round trip every time something works, something every time the user does something. And so typically for them,

Speaker: what the only solution is to bring more data to the client, right? So that you keep the interactions all on the client. And the problem with that is you end up with a sync problem, right? So if you bring data from the server to the client, well, the next question you're goingnna ask yourself is, is this data still fresh, right? And if you if you want the data to still to be kept up to date, then writing it in in Skip and using the Skip framework is is a really good fit because your data is always going to be kept up to date um naturally without you having to do anything about it.

Speaker: But the downside, and that's the trick with the the adoption of the framework, that's where it gets tricky, is that you actually have to rewrite the server-side logic. right So that's that's the problem, right? I mean, it's going to be easier when you start a project from scratch.

Speaker: But if if you have an existing backend with complex logic, it's going to be trickier to make it fit into the the framework. So is this something you could add on then? Is it something where if you had existing server-side logic that calculated something, had an API, would you then say, okay, we want a react reactive version of this, we'll have the back end of skip, connect to that API, pull it every three seconds for changes and turn that into a reactive broadcast mechanism? Yeah, you can do that. You can definitely do that. But

Speaker: So polling is natively supported in the framework, of course, and this is one way where you're going to turn things into something reactive very quickly. um You have to be very careful when you do that because I think ideally what you want is more like look at this API and see if you can make it itself also reactive. That is the best solution. Because if you're going to poll...

Speaker: you you are going to run into you know spike in CPU usage, um then you you have this horrible trade-off, which is that either you don't pull very often and then the data is not very fresh.

Speaker: and If you're trying to build an experience that feels a bit live, a bit collaborative, then this is really not nice. um Or you pull very often, but then you have a CPU usage trade-off problem. so I would say Yeah, ideally, um if you really want to build a real-time experience, a reactive experience, you would have the whole logic under a reactive setup. Because if you have that, then everything is just going to feel magic.

Speaker: You're just going to change something. Everything is going to update across multiple machines, and you won't have to do anything. But short of that, if you have to integrate into existing systems, then polling is it's probably the way to go.

Speaker: yeah Yeah, I can imagine very easily a system where you can't change it, but it's better to have one process polling every 500 milliseconds than 10,000 clients or polling every three.

Speaker: Yeah, there will be situations like that where you can basically share the polling. So having where polling is actually better because if something is very, very hot, then you can share the the compute they took to respond to you know a query.

Speaker: But it can bite you the other way, where you had a system that was barely used and very, very cheap called once a day, and you end up calling it every 100 milliseconds, and now you're not spending the same CPU.

Speaker: Yeah, yeah, yeah. Yeah, I guess you've got to analyze the existing system first, which is always good advice. taking Yes. Okay. Yeah. so is i'm I'm thinking, you may be finding this too, right now i am surrounded by sports yeah from Wimbledon to the World Cup.

Speaker: um I'm not a consumer myself, but I know people are very excited by it. And that's a system where not much happens until something very, very important happens and everyone wants to know instantly.

Speaker: that's a system where reactive broadcast to many clients would seem to be a very natural fit. Yeah. I mean, ah anything real time collaborative, which is definitely sports, ah live sports is definitely fits in that category, yeah is a natural fit.

Speaker: and So that makes me wonder, to what degree of update traffic and what degree of number of concurrent reacting clients have you benchmarked this with?

Speaker: So we have several versions, right? We have the version with um with ah the JavaScript version, which obviously is not as performant as the native version. Fair enough. And I used to know those numbers, but I don't remember the top of my head.

Speaker: um But if I'm not mistaken, and it also depends a lot on the size of your... um of the complexity of what what's you the the the logic is actually doing. So if, let's say, you have changes that come in, your logic is super, super simple. It's basically a routing problem. And really what you're measuring in that scenario is how many connections can node take. You're really basically measuring the the ability for

Speaker: ah for a node or a button to take traffic, right? So not super interesting. The moment you add CPU, the reason why it's difficult to give numbers is because... ah The moment you add logic, sorry, um it's difficult to give numbers because it's going to be very logic-specific, right? So if your logic is simple...

Speaker: then we will be almost as fast as you know if you were doing nothing. like If your reactive system is adding one to a flow of integers, I'm pretty sure that whatever number of connections that you're going to see and traffic that you're going to see is going to be very similar to something doing nothing, where you just you know rerouted the traffic. But now if you're you know your system does a lot of compute, then that will slow down dramatically. So I know it's a very frustrating answer, but basically um the benchmarks are very, very much dependent on

Speaker: what you want to compute and how complex that is. um Yeah, so that's fair. I can see plenty of situations where the logic was simple enough that it would be dominated more by serialization costs.

Speaker: Yeah, definitely. Definitely. it's not... um but it's not It's not super interesting to use the Skip framework if your logic is simple.

Speaker: Basically, if you have well identified... So here's the question I always ask people before engaging with Skip. I tell them, How hard would it be for you to write your logic backwards? Meaning i give you a change, right? I'm not asking you to write the logic. i'm I'm asking you, if I give you a change, how hard would it be for you to figure out what you have to update and and send and figure out to whom you should send the right messages, right?

Speaker: And if the answer is extremely simple, then you're probably dealing with a routing problem. And if you're dealing with a routing problem, you should probably not use skip. That's not what it's for. So typically imagine you have chat, right? You have a chat, you have a message that comes in.

Speaker: um How do you figure out who that message goes to? That is... probably simple enough that writing it in skip is going to be a huge overkill.

Speaker: You shouldn't do that. You should just write it in whatever language you want and it's just going to be fine. That is a routing problem, right? However, yeah if you find yourself with, you have a chat, but you're trying to maintain...

Speaker: I don't know, something, some property of the users of this chat, right? You're trying to maintain how many people are active and then of those who are active, who are friends with each other and those who are friends with each other, you want to the connected components of the friendship relationship that they have between between the people who are online, right?

Speaker: obviously fabricated a problem. Now we're in a very different realm. right like Now, if somebody disconnects, ask yourself, so do you know which connected components must be updated in your thing? And you'll find yourself scratching your head thinking, that's actually hard. And that's what I mean by running the logic backwards.

Speaker: If running the logic backwards, when I give you a change and off the top of your head, you're incapable of telling me what needs to be updated and how, then you are in a good place for for skip the framework. That's where it's going to shine.

Speaker: Yeah. So the closest I can think to that particular example is if we were playing a real-time strategy game, when i dis When I disconnect, all my troops disconnect from the universe.

Speaker: right so So that makes me wonder, do you have much traction in the kind of gaming world? So we don't for now, and I think it it really has a lot to do with the fact that it's a world that I don't know at all.

Speaker: so i yeah So I think it's it's it's a super interesting world, and I think we could do a lot there. But um yeah, not much traction in the gaming world, just because I don't know anybody. I am very...

Speaker: tooling language biased as you can imagine. So I guess that that's where my network is. But yeah, if you know people who'd be interested to work with us on games all year, I'm really happy to ah to engage with them.

Speaker: I'll put your contact details in the show notes, anyone listening. ah So I guess that bridges into a a couple of questions, ah commercial questions. So Skip is both the language and the framework and the company, right?

Speaker: So what's the licensing for the language and the framework and how's the company doing? So the licensing, we have decided to put all under the MIT license.

Speaker: We wanted it to be ah very open. um And the strategy for the company is more to build products that um are based on the technology that we have rather than sell the technology itself.

Speaker: So okay the latest product that we built is called Skipper. So Skipper is a coding agent that operates in a closed loop. So you can find it in skipperai.dev.

Speaker: What that means is that it's a coding agent that works a bit like a compiler. So you give a spec. And then you run Skipper, and Skipper is going to produce a working program out of this spec without um you in the loop at all.

Speaker: It will just iterate by itself. And that was a unique fit for reactive programming because what ends up happening is the the AI is...

Speaker: ah proposing changes, and you want to take these changes into account as quickly as possible. So that was a good fit for us. And we ended up building a ton of tooling to make that work. So we rewrote an implementation of TypeScript from scratch.

Speaker: that Really? Yeah, that is both sound and incremental. So we have our own TypeScript. um Oh, wow. The reason why we did that is we wanted it to be sound.

Speaker: Because if it's not sound, then all the other analysis that we wanted to do after, we would not have been able to do. Remind me the definition of sound. So sound means that you can trust the types.

Speaker: So if there is written int x, that means that at runtime, you know for a fact this is going to be an int, right? That's sound. While in an unsound system, such as Hack, such as TypeScript, such as many others, there is written number, but it could actually be null or a string at runtime. And you know the that there is written number there is just a best guess. right It's not a proof of anything.

Speaker: It's not baked into the runtime system? or it's not It's nothing to do with the runtime. It's a static property. So you have a type system and you have some type system like Haskell, which are correct. So sound basically means correct, which means that once the type checker says, okay, this program is correct, you now have a proof that at runtime, you're going to get the types that you expect, right?

Speaker: And you have some other type systems that are best effort. So they try their hardest to tell you what the types are going to be and enforce that.

Speaker: But there are many places where they're wrong. And so these are called unsound systems and TypeScript is unsound. So we needed to build a sound version of TypeScript to get all the other tooling that we wanted in place to build this. ah So this is a typical example of we are not trying to sell Skip the Framework, but we're using Skip the Framework to build a product.

Speaker: um and And that product is commercial product that you need to pay to use. And it's not under MIT license. Okay. okay So, Skipper.

Speaker: Give me a little bit more on that before we go. ah so okay well So, what's the vision? the vision is

Speaker: LLMs are going to write more and more code. And it's only a matter of time before there are some classes of projects that are entirely managed by LLMs.

Speaker: So where I go to write a service and from the beginning, from the the moment I'm thinking about this service until this this service is in production, everything, you know all the code, all the management, all the deployment is all done by an LLM.

Speaker: And if that's the case, then the differentiating factor for this kind of coding agent is going to be latency.

Speaker: so Imagine I'm writing, I have this kind of LLM that's writing this kind of program. What's going to happen is those programs are going to grow.

Speaker: right Right now, the coding agents, they're still working on relatively small program when they write them from scratch. But eventually, those code bases are going to grow. And what you will want is to do two things. One, you want to give feedback to the LLM very quickly on what it's doing, if it's working or not.

Speaker: And two, you want to avoid um drift. So you don't want an LLM that produces, that just add patches and patches and patches on whatever it's doing, because it's going to accumulate technical debt.

Speaker: And so what we wanted to do was create a development environment that is super constrained, where we really have everything under our control, and super fast. The way it gives feedback to the AI is is always instant.

Speaker: And so that the AI can iterate by itself very quickly on the code. and Or when you ask for a change. So you ask for a change, we can take that change into account between the moment where um we analyze the spec until it's live in production, under a second, the whole even on a large code base, the whole thing just updates very quickly. That's that's Skipper.

Speaker: Right, yes. I have definitely had the experience LLMs where while it's working on a change, I have more spec to send it and it becomes this asynchronous system which I wish was more reactive to my intent.

Speaker: That's the idea. The idea we, but the for that to work, you really need to be able to trust the outputs on the code. So the LLM is going to output code and you need to trust that output. And to trust that output, our approach is you need to constrain it and constrain it like very hard. Hence our own version of TypeScript.

Speaker: Okay, yeah, yeah. That leads me to one, I'll make this the last question, but in a system like that, what would you what you would also want is um an incremental test suite. Have you ever looked into like just recomputing the necessary tests to rerun?

Speaker: So yes, exactly. I mean, that's the whole idea. i mean, let's let's do it together. So let's say you take a diff. The diff comes in. The first thing you want is um the type checking to be very fast.

Speaker: So even if the code base is big, you don't want the changes to type check in tens of seconds. right So the type checker comes back, is incremental, is fast.

Speaker: You iterate like that until this is finished. yeah Once this is finished, because the type system is sound, you can actually trust the types. Since you can trust the type, you can run a reachability analysis. And the reachability analysis is going to tell you which parts of the code bases are affected by this change, which includes the tests. You will know which tests to run.

Speaker: because now you have that. So if your reachability analysis is also written in skip and also fast, now it will respond instantly and tell you, yes, you have to rerun that test, that test, and that test. Then you keep on iterating until you know the all the tests pass.

Speaker: right Once all the tests pass, If you have taught your LLM to use the skip framework, then what is actually generated, the program itself, is also incremental.

Speaker: So you don't need to restart everything from scratch. You just need to invalidate the code that needs to be updated and go recompute the pieces of state that needs to be updated.

Speaker: And so now you have the whole chain... that is incremental. And my eyes are sparkling just talking about this stuff. I see the future you're going for. Yeah, absolutely. Nice.

Speaker: Okay, that means I have to go and give this a try. um So actually, not all of this is plugged in. So I just don't want to give false expectations. Right now, we have a harness. We have some of the tools I mentioned are plugged in, but not all of them. So this whole completely incremental experience is not fully ready yet, but we will roll roll it out piece by piece.

Speaker: Okay. You have to promise me I'm going send you away now. I've kept you locked up in a hotel room in Paris on a sunny day for far too long. It's fine. You have to promise me you go and enjoy that rather than carrying with on with the framework right now. I will. I will. I'm due for, you know, beer and a good dinner. it will be perfect.

Speaker: Absolutely. Well, thank you very much taking me time talk to me because I really enjoyed that. And you have had a fascinating career to date. Thank you. Thanks a lot. It was very, very fun for me too. So anytime you want me back or anytime you want me to talk to a ah fellow language geeks, it's the crowd that I tend to do best with. and You know, hook me up.

Speaker: Brilliant. And next time you're in London, let me know. cool I'll let you know. Julian, thanks very much. Thanks. Thank you, Julian. i don't know if you could tell in that recording, but generally, as a rule of thumb, I like podcast episodes to be about an hour, but this is one where I just got to the hour and 20 minute mark and thought, no, we're going as long as this takes because I'm enjoying this and I want as much as I can get.

Speaker: I hope you felt the same way. If you did, please make sure you click like, subscribe, maybe share with a friend and stay tuned because we'll be back soon with another episode. Until then, I've been your host, Chris Jenkins. This has been Developer Voices with Julian Verlaget.

Speaker: Thanks for listening.

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Recommended