Featured image of post Apache Pulsar with Jowanza Joseph

Apache Pulsar with Jowanza Joseph

I talk with Jowanza Joseph about his new book Mastering Apache: Pulsar Cloud Native Event Streaming at Scale.

Listen

Subscribe

Transcript

[00:10] Tim: Welcome to episode 8 of the Into the Hopper podcast. My name is Tim Hopper, and I’m here today with my friend Jowanza Joseph. Jowanza is a software developer in Salt Lake City and, as of last year, Vice President of Software Engineering at Finicity, a Mastercard company. In December, Jowanza published what I think is his first book, Mastering Apache Pulsar: Cloud-Native Event Streaming at Scale, with O’Reilly Media, and we’re gonna talk about that book today. I almost said welcome, O’Reilly. Welcome, Jowanza.

[00:41] Jowanza: Thanks for having me on, Tim. I really appreciate it.

[00:44] Tim: My pleasure. This is a little overdue on my timeline, but we’re right after the release of the book, and it’s a perfect time for people to pick it up. Anything else you’d like to tell us about yourself by way of introduction?

[00:58] Jowanza: Yeah, I’ll add a few things. So, I’m an avid road cyclist. So, that’s something that I have really enjoyed after moving to Utah. So, I’ve been in Utah for, oh, it’s probably like 13 years now and something I picked up pretty early. And then, I’m a Father. So I have 2 kids. I got a 5-year-old and a 2-year-old and also a husband. And that’s about it. I think that covers my landscape of the landscape of Jowanza.

[01:31] Tim: I mean, model rocket builder.

[01:33] Jowanza: Oh yeah.

[01:34] Tim: Drone photographer.

[01:36] Jowanza: I don’t want to overstate my skills. I think before I met you, like I thought, oh, I’m a photographer. And then I saw your stuff and I’m like, I’m not a photographer. I’m just I’m just a guy with a camera that pretends because your stuff is amazing.

[01:51] Tim: You’re very kind. I’m trying to find time to photograph anything other than my kids, but that’s been pretty limited lately.

[02:00] Jowanza: Yeah, why can’t the day be like 35 hours instead of 24? And then we could have that time. But do you think our employers would take that extra 10 hours from us anyway if we got 10 more per day?

[02:13] Tim: Probably so.

[02:14] Jowanza: Yeah, probably.

[02:15] Tim: Maybe that 4-day work week.

[02:17] Jowanza: Yeah, no kidding.

[02:18] Tim: So, you wrote a book on Apache Pulsar, which I had never heard of until you started writing this book. What is Pulsar?

[02:27] Jowanza: Yeah, so, Pulsar, the way I would describe it, is a new way to think about storing and retrieving streams of data. And so, I think most people would probably approach it with the definition that’s like, It’s similar to Kafka, but has these key differences, but I really look at it more as a, hey, if you were to think about streaming data, and how a whole topology of streaming data would come together, how you would store it, how you would retrieve it, that’s really what Pulsar is designed to solve, is that kind of bunch of problems, and then the use cases for it come as as more of a kind of applying what Pulsar solves to your use case, but it’s really solving kind of a very specific thing, I feel like.

[03:17] Tim: How did you get into it? I guess that leads into the question of what motivated you to write a book on it.

[03:24] Jowanza: Yeah, so let me talk about how I got into it. So the first time I used Pulsar, it’s that kind of right shortly after it was open sourced. At the time, I was using Kinesis, Amazon Kinesis, for some of kind of the same workloads. And I was very compelled by some of the key differences in Pulsar that we’ll probably get to in a minute about, you know, the— well, one, the protocol, and then two, kind of the client interactions with it. And so I, at that time, migrated some of our workloads at the company I was working at from Kinesis to Pulsar, which was kind of crazy because it was like it had been out for like a year, like in public. But I guess I had a lot of leverage to do that at the time. And so, that’s when I first got into it. And then subsequently, I left that company. I didn’t get to work with Pulsar anymore, kind of on a work basis.

And then, actually, with this book in particular, O’Reilly approached me about it. And so, I’m not particularly sure why. I’m not a committer on Pulsar, I’m not, you know, I’m actually, at least at the time, wasn’t using it in production or anything like that. But I think a lot of the past writing and speaking I did about Pulsar when I was using it, kind of, that’s kind of identified as someone that would be a reasonable target to write it. And so they approached me, I wrote the proposal, and they accepted the proposal, and then the rest is history.

[04:55] Tim: People can find some talks I saw that you’ve given on YouTube on the topic, which is probably a nice introduction. And this wasn’t that long ago. I mean, I think I read it was open source only in 2016. So, um, yeah, less than 6 years ago, probably.

[05:10] Jowanza: Yep. Um, but yeah, it was, it was, uh, created in Yahoo, you know, a lot, a lot earlier than that. But yeah, it wasn’t open sourced until, until then. So, you know, I think it was pretty, pretty hidden, um, as far as like anyone knowing or hearing about it, probably until like 2018 is when you saw a lot more marketing, and use cases, and blog posts being written about it, but it’s a lot smaller audience, I think, than something like Kafka, which has been able to benefit from, well, one, the timing, and then two, just the much wider adoption, and being a little bit earlier, too, to being open-sourced.

[05:53] Tim: That makes sense. And I saw 2018 is when it became a top-level Apache project, which certainly gives it a higher profile.

[06:00] Jowanza: Yep.

[06:01] Tim: So, who is your target audience for this book?

[06:03] Jowanza: Yeah, so, that’s actually an interesting question, because I’m gonna see if I contradict what I wrote as the target audience in the book. Yeah, so, the target audience is, I think, I kind of approached it from 2 angles. So, one is, you know, people who are interested in streaming, and, like, stream processing and streaming in general, it’s kind of like an architectural paradigm. because the book focuses on building up from the idea of why we need this type of system and interaction. And then also people who are using Pulsar, so people who are getting started with Pulsar, to act as a reference guide for them of all the operational procedures, how things work, how to construct different elements of the Pulsar ecosystem. So that was really the two-prong audience. And I think for either group, you can start with the book and kind of get value out of it over time. Hopefully, that was the goal, at least.

[07:03] Tim: You know, when I heard of this topic, I’ve done some streaming stuff in the past, mostly in Kafka, and have an interest in kind of distributed systems and things, but it hasn’t been as much my priority in terms of kind of on the management and backend side of building these types of systems. So I didn’t necessarily think it would be a book that would be right up my alley, but— and I actually haven’t read all of it, but I read a good chunk of it in preparation for this, and it was very engaging and had a lot of— even just some of the foundational material in the first few chapters and thinking about event processing really helped solidify for me some concepts that I’ve experienced in the workplace but never really studied. So I think Probably more people might benefit from this more than they think if they’re dealing with streaming data and event processing. I think you lay out some things really nicely.

[08:05] Jowanza: Oh, thank you.

[08:06] Tim: Actually, on that note, I wanted to ask: in the first chapter, you have this great story of selling Pokémon cards in school and use it as an analogy for distributing events and information through a system. The chapter’s called The Value of Real-Time Messaging, but I wanted to know if that was a true story, because it was pretty great.

[08:32] Jowanza: Yeah, thanks. So, in kind of going back and reading back through, because I got a lot of positive feedback from the editors on that, what I kind of realized is actually it’s a combination of true stories, but the end-to-end story itself wasn’t necessarily how it happened, right? So we did have a system for communicating via walkie-talkie, but it wasn’t actually for the Pokémon cards. That was for something else.

[09:00] Tim: Okay.

[09:01] Jowanza: But we did sell Pokémon cards. And so I think my imagination and memories from whenever 4th grade and 5th grade kind of merged back together to make this story. But then when I kind of really drilled in on like, okay, who would’ve been involved in this? Like, how would we have done this? I kind of realized, oh, this is actually 2 stories that I’ve merged into one. And so I think there’s not anything that’s like false about it, but it’s really not end-to-end like the one story that happened. So it’s really 2 different situations, but you know, it works.

[09:36] Tim: Yeah, and I really liked it. I thought it was a very engaging way to start the book. Had some helpful ways to think about the problems.

[09:44] Jowanza: Thank you.

[09:46] Tim: Yeah, I was pleased. I mean, you know, you hear a topic like this, and you kind of expect to just jump into something very dry and technical right away. And you did a nice job not doing that. So I appreciated that. You must have been very cool running around with walkie-talkies on your little side hustles at school.

[10:05] Jowanza: Yeah, it’s actually kind of interesting. So I know this podcast isn’t about this topic, but I grew up in Queens, New York. And so we— and then I grew up kind of in a lower-income neighborhood and we didn’t have walkie-talkies that our parents bought us. So we actually found some in a kind of like a trash situation, like open trash bags. And so that’s kind of how things sort of originated with that. So it’s a So we probably did seem pretty cool and futuristic, but— or not futuristic is not the right word, but just kind of cool because no one else had walkie-talkies. And so that must have been really kind of interesting to everyone. But yeah, it was a good time.

[10:54] Tim: Kids these days are walking around using their Apple Watches in walkie-talkie mode, probably.

[10:59] Jowanza: Yeah, my daughter, she has like an iPad. And, and, and so I’ve tweeted about this where you, the auto switch on the AirPods, right? I walk into my house and then I’m listening to whatever she’s watching and I think, well, I didn’t have an iPad until I was in college. Like, and she’s like, you know, just chilling on the couch with, you know, iPad. It’s crazy. ‘Cause all the cost of all this stuff has come down and it’s like, just kind of like normalized now, which is interesting.

[11:28] Tim: Yeah. It’s a different world. So kind of at a high level, the book is called Cloud Native Event Streaming at Scale. Can you maybe tell us what event streaming is as a concept, and then why you might use event streaming versus something like just a REST API for systems to communicate? We often think about these days microservices communicating through REST APIs. REST or gRPC or something along those lines? Where does event streaming differ from those?

[12:03] Jowanza: Yeah, so, I would say, let me describe event streaming. Event streaming is kind of, instead of being— actually, let me try to contrast it. That’s the easiest way to explain it. You kind of said, oh, why would you do this versus REST? In the REST world, REST is all about resources. your APIs are built around operations relative to resources, so creating new things, editing them, deleting them, and, like, oh, so creating, reading, updating, and deleting them, right? And so it’s not about keeping track of the full state of those objects or resources. So you can have a REST API that’s about like creating new pictures or creating pictures. You could create a new picture, retrieve one, update one, delete one. But if you wanted to understand the full history of that picture, it’s usually not something that would be on that same resource API.

Maybe there’s another one that will give you a full sequence of everything that happened. Contrast that with events. It’s all about keeping track of the full state of events in a resource or events in general. So, it could be kind of abstracted away from a resource. If you look at the same picture idea as events, it’s about, okay, so, creating the picture is one type of event, modifying the picture is another type of event, deleting the picture is the third type of event. And the idea is that you, as a consumer of those events, can then reconstruct all the history of that event or get the exact state of the event right now in time, but from one kind of specific consumption of a topic or a queue. So, it’s a paradigm shift to think about everything as events rather than as a resource.

And so, then you get a lot of different types of interactions that you can have as a result. And so, why, I guess, I didn’t answer the, like, Why would you want to do one versus the other? And I think I talk about this a little bit in the book. It’s modeling kind of a resource or, like, kind of database-oriented interaction in an event world is actually possible, and not that it’s trivial, but you can kind of recreate the same kind of interaction from, like, a stream to a table, but going from kind of a table, like, back to a stream is more difficult because if you think about how keeping the state of a row in a database or an ID in a database is much— it’s really kind of not what databases are for.

It’s kind of to give you a mutable copy of, like, okay, we’re just going to keep mutating the same row, whereas the event paradigm is we’re going to just always append and not mutate anything that’s existing already. So, it’s a very contrasting view. Anyway, that was a long answer, but that’s my answer.

[15:25] Tim: That’s really helpful. And in the book, you say, at its core, Pulsar’s implementation is a distributed log. And you talk about the 2013 blog post by Jay Kreps, who was at LinkedIn at the time, called The Log: What Every Software Engineer Should Know About Real-Time Data’s Unifying Abstraction, which I actually started Right that same month, I started working my first job, kind of dealing with distributed systems and log-based architectures. And I read that blog post just over and over, like, I need to really understand these concepts. But I had not really connected mentally until reading this about kind of an event system being an implementation of a distributed log, which is what you call it, which I think is a pretty fascinating— and it’s sort of a simple concept, but it’s a fascinating concept, I think.

[16:22] Jowanza: Yeah, I agree. You know, the distributed log is— so since the creation of, like, or the open sourcing of Pulsar and Kafka, Other messaging protocols have appended a distributed log to their messaging layer to give them the same semantics. RabbitMQ has RabbitMQ Streams now, and then Redis has Redis Streams, and then NATS has this thing called Jet Streams. And essentially, yeah, it’s just kind of— I’m not trying to oversimplify what they’ve engineered, but it’s really, okay, how can we take this queue idea And instead of making it, you know, sort of not, like, not really— it is append-only, but it’s also— well, it’s not append-only, right, in the queue, right? It’s append and also, like, delete. And so, how can we turn that into, like, append-only? And so, yeah, it’s an important distinguisher, or, like, yeah, distinction between kind of what, like, Kafka or Pulsar is trying to do, and then what, like, a message queue is doing, and how that leads to different types of interaction paradigms and scalability paradigms as well.

[17:35] Tim: That makes sense. I asked some audience of the podcast if they had any questions, and one wanted to know kind of where Pulsar compared to Redpanda, which if you can maybe explain a little bit what Redpanda is, I’m not actually familiar with it.

[17:50] Jowanza: Yeah, for sure. So they’re very different. But so Redpanda is taking the Kafka protocol. So Kafka has a protocol that says like, this is how you have consumers and producers, and these are the types of things they can do, these are the types of configurations they can have, and then this is the behavior that they should have. And then instead of having that implemented in Java, they reimplemented it in C++. So, you still get the same Kafka API and protocol, but just internally, it’s actually different. And so, what you get is, better memory management, and also the GC issues that you can have with Kafka at scale, so garbage collection, excuse me, kind of go away because they’ve just reimplemented it in C++. That’s what Redpanda is.

Pulsar, by contrast, is a completely different protocol than the Kafka protocol, so you can’t take a Kafka client and connect it directly to Pulsar. That said, Pulsar does have a framework for doing other protocols on top of Pulsar, so then Pulsar kind of becomes similar to Redpanda. So you can do Kafka on top of Pulsar, so there’s a kind of a thin layer of a Kafka protocol, and then it’s translated down into the Pulsar protocol and then back out. So In that way, they can be similar. So, there are some companies out there who sell Pulsar as a service, and then they also offer Kafka on top of Pulsar as a service. And the idea is that it helps you migrate over to Pulsar. So, you kind of start doing Kafka on Pulsar, and then maybe you start kind of doing Pulsar alone. And so, that is— yeah, that’s one of the implementations.

And actually, that same core idea of, like, taking the protocol, and then kind of reimplementing the internals of it, is something that Pulsar, kind of, as a community, is interested in. So, what they’re doing is just actually trying to modularize everything internally, so that if they wanted to just kind of, like, reimplement, kind of, the core Pulsar broker in C++ or Rust or whatever, that they could do that easily, whereas, you know, in this Redpanda project, it’s really like a complete fork of Kafka, and not like a rewriting of the internals of Kafka. Okay.

[20:26] Tim: So if we can maybe take a little step back there, I think a lot of people are familiar with Kafka, probably maybe secondarily Kinesis these days ‘cause of AWS. And you’re talking about Kafka is very different from Pulsar in terms of protocol, But at least my understanding is sort of similar in terms of the functionality they’re providing. Where would someone decide to use one or the other if you’re implementing it? Or maybe even when you initially implemented Pulsar several years ago, why was Pulsar a better alternative to Kinesis at that time?

[21:10] Jowanza: Yeah. You asked a loaded question. So, there’s a few reasons. So, let me— I think one important distinction to make is kind of going back to the initial question of, like, what Pulsar is, and I used the word storage. And so, one of the big differentiators between Pulsar and Kafka is that, and, you know, really any other kind of comparable technology, is that Pulsar the process of kind of reading and writing messages, so communication to consumers, is separated from the storage of the data. So the brokers that you interact with on Pulsar, so you would, if you wrote a client, you would connect to a broker, those are actually stateless, and then the data that’s in the topic is stored in BookKeeper. And so what that gives you is you can scale out the data, like, as kind of wide as you want to, and keep kind of a thin broker layer, right?

You can keep it as small as you need to meet your requirements around I/O. Whereas with Kafka, the storage and the computing are actually together, right? When you deploy Kafka, especially in the most recent versions where they’ve removed ZooKeeper, everything happens on the broker. So if you want to add more storage, you need to add a new broker. Kafka, I think, was initially presented as foolproof storage—you could just store messages in it forever—but what people started finding at scale is that that’s actually kind of hard to do, given the architecture.

So, with Pulsar, it’s a lot easier to do, given the architecture, so that would be one reason why you might want to look to use Pulsar is, hey, I want to actually use my messaging or my stream system as— I hate to use the phrase as a database, but like infinite storage. I want to store everything in here. I always want to have this available. Another big thing with the difference is multi-tenancy. Pulsar supports multi-tenancy out of the box. If you want to really logically separate workloads and also limit workloads by resources. So, hey, for this namespace, they should only get this much CPU, and their messages should have this much throughput. That’s already, like, a native concept, and it’s not in Kafka. And then in Kinesis, it’s— I’m actually not sure, I don’t remember if Kinesis has namespaces or not.

I would think no, but those are kind of the 2 biggest things. And the 3rd one is kind of a smaller one where Pulsar, it’s kind of the base implementation of their topics are actually a single partition. And so what you get with that is kind of a more familiar system to like a queue where if you use RabbitMQ or something like that, those would be like a single partition topic, right?

[24:16] Tim: All—

[24:17] Jowanza: every message goes to kind of one queue and then it gets of fanned out, or it goes to one consumer on the other end. And so Pulsar can support, like, that, you know, single partition and multi-partition. So if you wanted to model something as a queue, or as a stream, you can do that, whereas you can’t quite as easily in Kafka, because every topic is a multi-partition topic. And so you get kind of different outcomes there. So those are— that was a really long answer, but there’s some— it’s very, like, nuanced reasons. And so I think that’s why for the adopters of Pulsar, they usually have some, like, real, you know, sophisticated thoughts behind how they want to do streaming and, like, how they want to do messaging at scale.

And I think for kind of the median, you know, to, like, even, like, the 80th percentile, you know, looking at this, it’s like, well, Kafka kind of fills all my needs, right? Or RabbitMQ does. I don’t even want to do streaming. And so it’s really kind of niched into, or a niche in, like, a really small kind of space of people who really understand, like, the limitations of the models that, you know, Kafka or other messaging services provide.

And so, I think that the kind of commercialization efforts around Pulsars, the companies that are doing that, are just trying to, like, make it simpler, and, like, not about these nuances and much more about just kind of providing an end-to-end platform, because Pulsar has some other features too I didn’t mention in that list, but that’s roughly, I think, the real kind of message layer differences are those I described.

[26:02] Tim: Yeah, that makes sense. Another question from the audience, are there any of the hosted Pulsar solutions that you particularly like that you know about?

[26:12] Jowanza: Yeah, so I will caveat, I’m not getting paid by anyone to say any of this. So, but yeah, I think so. Stream Native, you know, they’re really the kind of pioneering and most— and like the largest player in this space. They have a nice managed solution. And for most people, if you want to get started with like Pulsar for free, you can use that and then just see if you like Pulsar. And it’s a pretty nice service, and you won’t get charged anything for like a small distribution. DataStax is another company that has something called Astra Streaming, and Astra is, you know, built on, on Pulsar. And similarly, like, if you want to just get started on like a pay-per-usage basis, they have like a tier for that. And it’s also, you know, very simple to get started with.

And one of the nice things about both StreamNative and Astra is they kind of force you to get started with good practices. You have to use authentication and encryption, right? They make sure that you get started that way, whereas some managed services will just be like, oh, here’s your password, go at it. But these tools force you into a good pattern, which I think is a good thing when you’re getting started. You don’t have to encrypt data or use good role-based access control, but these tools will force you to do it.

[27:47] Tim: Okay, very nice. One thing in the title of the book, you talk about Pulsar being cloud-native event streaming. What does it mean for the tool to be cloud-native?

[27:58] Jowanza: Yeah. I actually hate the term cloud-native, so it’s funny that you asked. But I think what’s meant by cloud-native, and so, also, I didn’t, you know, fully decide the title there. So, I think that I definitely approved of it and said, great. But, like, I think if it were me, you know, I probably would have called it something else, maybe. But—

[28:25] Tim: As long as it sells books.

[28:26] Jowanza: Yeah, right, right. So, I think what cloud-native has come to mean is, like, it’s sort of a CNCF, like, a Cloud Native Computing Foundation term, right? That is sort of like, my project works with CNCF things. And so, CNCF includes Kubernetes, right? That’s the biggest project out of the CNCF. But then, it’s also including things like Prometheus, right? So, if you want to store your telemetry data, right? You have a Prometheus exporter. And then it includes things like— I’m trying to think of some other big projects. So, they include, like, you know, other container image formatting, or, like, I’m trying to think of what you would call, like, some of their container things. Anyway, so, I think what it’s come down to is, like, cloud-native really means it runs on Kubernetes.

This is kind of how I interpret what people mean when I hear the word cloud-native. It’s like, this will just run as a Helm chart, or, like, this will run as an operator, or, like, you know, this is embedded in Kubernetes. Because I don’t think it means, like, hey, you can get some EC2 instances and, like, figure out a way to put this on the cloud, right? I think that it It means something more specific that’s like, this will run on Kubernetes. That’s how I interpret it.

[29:50] Tim: Yeah.

[29:51] Jowanza: Pulsar, as an open-source project, is very invested in Kubernetes. It’s like, if you go to how to install Pulsar, you’ll look through, and they’ll show you how to do it on bare metal. They’ll show you how to do it on EC2 instances. But then, when you get to the Kubernetes one, you’re like, oh, okay, this one has a lot more automation and simplicity, and all of the attempts to make installing it better really kind of point back to Kubernetes. Then the extension features in Pulsar that we haven’t talked about yet, those are all trying to run better on Kubernetes. I think when I hear cloud-native, I hear Kubernetes. Maybe that’s cynical, but that’s how I feel about it.

[30:39] Tim: If you don’t like that, direct your email to Jowanza, not to me. Maybe if you want to tell us what are Pulsar extensions.

[30:47] Jowanza: Yeah, so I think one of the underrated parts of Pulsar is actually the ecosystem. So I’ll talk about a couple of them. So one is Pulsar Functions. So Pulsar Functions are kind of a lightweight compute execution engine for Pulsar that allows you to use a Pulsar topic as an entry point or like a source, and then a Pulsar topic as a sink, and then you can do logic within the function. So you could think of it as like a— I don’t like to use the word stream processing because I feel like that’s an overloaded term, but it’s really just a compute paradigm where you have Pulsar topic as input and output. You can think of it similarly to Lambda functions where you can have events as an input, and then you can— but you can have more than events as output in Lambda functions. You can do a lot of things, but that’s a base concept.

So Pulsar functions generally are supposed to be lightweight and supposed to kind of be a small domain that each function is trying to do. But there is a concept of state that you can establish among the functions. So there’s a centralized database in ZooKeeper, but— or not ZooKeeper, BookKeeper, excuse me. And so, you can do some pretty sophisticated things with, like, global counters and, like, global state management with Pulsar Functions. And then, Pulsar Functions, you can write them in Java, which is kind of the core language of Pulsar. You can write them in Python, and you can write them in Go programming language as well. The other Pulsar— Oh, go ahead.

[32:28] Tim: Well, do functions come— I mean, that comes with Pulsar? So if you install it, you’re ready to do that?

[32:35] Jowanza: Yeah, yeah. So with a default distribution of Pulsar, you can run Pulsar functions and they will run on the same brokers. And so you can imagine there’s some limitations to doing that at scale. And so if you’re really running Pulsar functions, you want to run them, have a separate deployment. And so you can do that. on Kubernetes cloud natively, or you can run them on VMs or whatever you want to do. But yeah, if you just installed Pulsar on your Mac, on, like, you know, kind or whatever, Kubernetes would be the easiest way to do it, then you would have Pulsar Functions that you could deploy there as well. Okay, so the other 2 things are— I’ll say 3 things. So the other one is pulsar.io. So you could think of Pulsar IO as analogous to Kafka Connect.

So it’s a framework that is, you know, trying to connect Pulsar to other elements of ecosystem, usually for the purposes of change data capture. So taking changes from MySQL, for example, and then writing those to a Pulsar topic. or taking events that are getting published to a Pulsar topic and publishing them out to Elasticsearch. So, it’s really kind of this framework to just kind of write once a good connector that will just connect 2 things, and it’s usually, you know, Pulsar is going to be on one end of those, right? A Pulsar topic either will be a sink or source, but then it can go, you know, reaching out to a database or, you know, a message queue, another type of message queue. So you might want to go have Pulsar topic in, RabbitMQ topic out.

And so the whole concept is a framework for building that type of connection. One of the interesting things is that Pulsar Functions and Pulsar IO share the base implementation. So when you make a Pulsar IO connector, it’s really like you’re building a kind of Pulsar function for executing a specific type of task.

[34:52] Tim: That makes sense.

[34:53] Jowanza: But yeah, you can only write them in Java right now. It’s kind of one of the limitations of it. So if you don’t like Java, then tough luck, I guess, with this, that for now. So, and then, okay, so then the third one is Pulsar SQL. So Pulsar SQL I would say it’s not analogous to ksqlDB. It has some similarities, but they really have different goals. Pulsar SQL is a way to read Pulsar topics in and then from their schemas and then write SQL queries against them. They are streaming queries similar to ksqlDB, but the goal isn’t to, like, use, like, Pulsar SQL as, like, a backend database for an application read, or it’s more for interactive queries is what it’s kind of there for.

So if you wanted to, you know, if you wanted to, you know, search over all of your transactions and, like, do some analytical queries on them just straight up from the topic, you could do that. If you wanted to join 2 topics together, you could do that. And the way that it’s implemented is it’s It uses Trino, which is formerly Presto, and it uses the metadata from the topic that’s stored in— or the topics that are stored in ZooKeeper, and then it can reach out to BookKeeper. Also, if your data is in tiered storage, it can reach out to the tiered storage destination, which I haven’t talked about yet. That’s the 4th and last one, is Pulsar supports tiered storage. What tiered storage is, is a way to automatically offload data from BookKeeper, where data is stored for Pulsar topics, into object storage.

You can actually also do it in HDFS or file storage, but that’s a weird thing to do, so I’m just going to talk about object storage. It makes a little bit more sense. The idea would be, Tim has a massive BookKeeper cluster, 27 petabytes of data, and his boss is upset because it’s costing a lot to keep that all. It’s like half the cost to put it in S3. So tiered storage would just have a kind of an automatic trigger based on either, you know, the length of the data being in Bookkeeper or, you know, the size of a topic, and then it would just write that data over into S3. And then the Pulsar brokers, which are the things that are retrieving data from Bookkeeper, you know, have an implementation that could kind of transparently retrieve it from the S3, like, if it needed to, like, for older storage that’s in the object store, right?

It retrieves it similarly to how it would retrieve BookKeeper. So as a consumer, you don’t really need to know where the data is stored. It just manages it for you. That’s the Pulsar ecosystem, roughly.

[37:51] Tim: That’s, that’s very neat. Are there interesting directions of current development, things that Pulsar is looking to offer in the future?

[37:59] Jowanza: Yeah, for sure. I didn’t get into this, but Pulsar uses Apache ZooKeeper for a couple different things. For one, Apache BookKeeper requires ZooKeeper to work, but you probably feel like, man, Apache projects need to name their stuff differently. ZooKeeper is probably familiar to you as someone who’s worked with Kafka or any other distributed thing in the Apache ecosystem.

[38:31] Tim: The main thing I know about ZooKeeper, and I assume you’re headed this direction, is it’s something that every project is trying to remove as a dependency.

[38:37] Jowanza: Yeah, and for good reasons. But, yeah, I think Confluent had a great blog post about, like, how, when they announced that they were gonna remove it from Kafka, everyone was kind of, like, dogging on it, and they were just like, hey, let’s defend ZooKeeper. This thing is amazing. It got us here. It’s scaling up to these ridiculous amount of events, but we’re just moving on to new paradigms, and we wanna do different things, and it doesn’t fit into the new things we wanna do. So, it’s not that it’s bad, it just wasn’t built for this next age of computing that we’re doing, and so we’re moving off of it.

So, Pulsar is removing that from— or, they’re not removing it, they’re making it pluggable, so that the use case that ZooKeeper serves in the Pulsar ecosystem is, you know, one, for BookKeeper, but then, two, to store metadata around each topic, and then how that data for the topic translates to what’s in BookKeeper. And so they’ve just made it basically an interface now. So if you wanted to implement your own way of doing that kind of metadata, you could. Or if you wanted to do like an in-Pulsar broker implementation, which is kind of a little antithetical to the whole way that it was designed, you could do that. Or if you wanted to switch it to etcd, you could do that. Or if you wanted to switch it to Redis or something else to store that metadata, you could do that.

So their approach is instead of being kind of more authoritative about it, but they want to make it more pluggable and different options for different use cases, especially on the scalability side. Each of those choices will have some trade-offs. That’s one of the fun things that’s happening. Another big thing is transactions, which are pretty new as of a couple releases ago, which essentially will allow an atomic operation of data being read into a topic, transformed, written into another topic, and so transactions are helpful, right, if you want to ensure kind of exactly once processing in a stream processing pipeline, but as you can imagine, orchestrating that is difficult, so Pulsar has an implementation in for it now. They’re trying to make it a little bit simpler and a little bit more rounded around the edges.

I’d say the third thing that’s interesting is Pulsar function mesh. So I mentioned before that, you know, the topology of Pulsar functions are pretty simple, right? They’re, you know, topic in, topic out. And what they’re finding is that, like, in the actual usage, people are trying to construct some, like, really elaborate function topologies. And so instead of, like, people abusing it, there’s just a framework in there now that you can construct topologies that may have a different source than Pulsar topics, or that might need to do multiple things as part of one operation in Pulsar Functions. Those are, I think, the most exciting things right now. Then there’s a bunch of other little improvements around the REST API and the admin API and things like that.

[41:53] Tim: And we alluded to this, but Pulsar is an open source project. So when we talk about Pulsar is doing these things, it’s open source contributors. Is there a company that’s primarily driving this and paying people to work on it?

[42:06] Jowanza: Yeah, so I think there’s 2 really big ones, right? So StreamNative I mentioned, and then DataStax right there. They’re both very good at kind of differentiating what should be proprietary and then what should be kind of contributed back. And so the things that are proprietary are usually the enterprisey features around, like, you know, role-based access control and, you know, different ways of doing auth. But the stuff like the kind of core capabilities or, like, enabling different types of, you know, exchanges for the metadata, that’s all open source. And so they’re both very active in the community and There’s other big contributors as well. So, like Splunk, they’re a big user of Pulsar, and so they contribute a lot of open source as well.

I don’t know if Yahoo is called Yahoo anymore, but they were the original creators, and I still see contributions. I’ve tried to make some of my own, but it’s one of these things where I’ll find a bug, and then I’ll start a merge require pull request, depending on if you use GitLab or GitHub. And, like, by the time I’m done, I’ll go back and, like, someone else has done it already. And so it’s happened to me 3 different times. So it’s a very active community, and, like, people are kind of thinking about the same problems and experiencing the same ones. So I’m hoping, like, it’s on my list of things to get, like, a nice contribution in there, but I’ve been beat to the punch more than once.

[43:41] Tim: So yeah, I mean, that’s discouraging as a contributor, but that’s encouraging as a if you’re using the project, like, people are really working on it.

[43:49] Jowanza: Yeah.

[43:50] Tim: In the last few minutes here, I’d be interested in hearing, what’s the process for writing a book like this?

[43:55] Jowanza: So, my process, I was very methodical to start. So, I had a— O’Reilly, I’m gonna give them a huge amount of praise for this. They’ve done this a million times, and so they have a nice system of, like, you get assigned an editor, And then they put a schedule together for you, and then they check in regularly with you if you’re making that schedule. And so, with the schedule they provided me, I went ahead and then put that into my calendar of like, okay, if I want to have chapter 1 done by 2 weeks from now, I need to write at these times, and I need to make sure that that fits into my lifestyle. Like, I don’t, you know, plan a vacation at this time, or like, I don’t— or if I do plan a vacation, I compensate for it by writing before. So I did that and I started out really great.

And then what happened is once I got into the more technical chapters that required me to take a step back and be able to explain the technical concept without being overly pedantic or dry, then I kind of lost steam. It took a while for me to say, okay, how am I going to talk about the bookkeeper quorum without just completely killing everyone’s attention. And so I slowed down a bit and then I picked it back up. And then I had a real rough time toward the end of the book. So the book was due in September, like the final copy. And then I lost my mom in August. And so it was like— I actually ended up being late because basically from the time that my mom passed away, probably until a month after that, I just couldn’t write anything. I was just too distracted.

[45:39] Tim: Yeah.

[45:39] Jowanza: And so then I picked back up, and like it was kind of a mad dash there at the end because I was like a month behind and we still wanted to make the date. And so I think the process is just really plan, and then like for me, I needed a lot of uninterrupted time, right? So if I had 6 hours on a Saturday morning, I would write, you know, 50 pages at that point of like, pre-page material, which would end up being like 10 pages in the actual book. But nonetheless, that was my process, is just like giving myself a lot of uninterrupted time and then planning out what I was going to do. And that’s kind of my personality. I know a lot of people take different approaches, but for me, it’s really kind of about planning and executing. And that works good for me, giving myself deadlines, like artificial or not, is really effective.

[46:31] Tim: So, we should thank your wife, Bethany, for giving you those Saturday mornings. And Bethany also illustrated the book, right?

[46:41] Jowanza: Yes, she did. So, she’s been my illustrator for my blog when I used to be a pretty regular blogger for many years, I think since we got married. And so, yeah, she did all of those pictures, which there’s a lot. There’s hundreds of them, I believe, in the book. So, she So I guess the process might be interesting to the listeners. So usually what happens is like I’m writing, I think of the idea, and then I draw a really terrible version of what Bethany eventually will kind of illustrate. So I’ll say, okay, I need something that’s like my Pokémon one. Like I need cards and like kids, and then she’ll just take it and make it. So she’s— yeah, I mean, I owe her a lot because I think some of the most positive feedback I got was around the illustration. So yeah, she’s a lifesaver and she made a huge difference.

[47:39] Tim: Yeah, awesome. And the illustrations are really nice, which makes it for a nice read that way. I have the Kindle version of the book. I think O’Reilly doesn’t sell books directly anymore. they kind of refer you to Amazon, but you can buy it in print and on Kindle. And the Kindle rendering is really nice, which, you know, that’s probably a plus of going with writing a book with O’Reilly. They really know how to do that because technical books don’t always translate well. Do you think you would write any more books? Yeah.

[48:10] Jowanza: So I think I would. I think if you were to ask me that question a month ago, I probably would say no. I think with some more time, I have a lot of— I’ve learned that I really like writing, and I like the process of explaining things through writing. And so, I would write another technical book, but I also think I’m interested in writing a memoir at some point too. And so, I think I need to maybe do something more important than I’ve done so far. But once I do, do something noteworthy for the whole world to see, then I think I will write a memoir to tell my story kind of from the beginning till that point or something. But yeah, I think writing is great. And I actually owe myself like 10 blog posts or so that I’m supposed to be writing now to like kind of promote the book. And so you should be seeing some of those coming out soon. I’ll post them on my Twitter. I’ve got like 10 drafts I need to finish.

[49:09] Tim: Very nice. Well, you’re one of my most interesting friends, so I’ll look forward to that memoir. People can go back and look. You’ve written on a variety of things on your website, jowanza.com, J-O-W-A-N-Z-A dot com. And I’d encourage people to go read some of your archives there.

[49:27] Jowanza: Yeah. Thanks, Tim. Appreciate that.

[49:30] Tim: And you can find Jowanza on Twitter, twitter.com/jowanza. I’m on Twitter, twitter.com/tdhopper, and I’m tdhopper.com. Anything else you’d like to share as we close?

[49:41] Jowanza: Yeah, I think the last thing is just, I think Pulsar and Kafka often get compared because they kind of seem like they take over the same space, and I even use them as analogies, but I just want to be very clear that, like, I consider them as very different, and that they solve some of the similar problems with a different approach. And both are really awesome and have great communities. And I’ve benefited from both, using both of them in my career. So I would not ever use the comparison point as like a disparagement point. I feel like it’s an important disclaimer. And one that I try to make is that they really have some really great advantages and disadvantages. And I just want to say that because I know sometimes things get twisted, but that’s my strong feeling is like, it’s a good comparison point because it’s easy to see. But in reality, it’s like they, you know, solve similar problems in different ways and they’re both great. And that’s it.

[50:40] Tim: Excellent. Thank you for coming on Into the Hopper.

[50:43] Jowanza: Thank you, Tim.

Feedback