Listen
Links
- Willem’s Twitter
- Adam’s Twitter
- Feast: feature store for Machine Learning (2020 talk)
- Feast
- Tecton
- GoJek on Wikipedia
- featurestore.org
Subscribe
Transcript
[00:10] Tim: Welcome to episode 5 of the Into the Hopper podcast. I’m Tim Hopper and co-hosting today with my former podcast guest, Adam Laiacano. Our guest today is Willem Pienaar, lead developer for the open-source Feast feature store library. as well as a developer at Tecton, which is a commercial feature store product. Willem and Adam, welcome.
[00:33] Willem: Hey, Tim, thanks for having me.
[00:35] Adam: Hi, thanks for inviting me back.
[00:37] Tim: Yes, sir, thank you very much. Willem, you recently changed jobs to work at Tecton, but you have a background in Gojek, which is an interesting company that probably a lot of people don’t know about. So maybe you can introduce yourself and explain your background a little bit.
[00:54] Willem: Yeah, so I’m actually, I’ll give you like a quick whirlwind tour. I’m a South African. I was kind of in technology for a long time, built a startup in South Africa around networking, worked in industrial automation and kind of ERP and kind of data warehousing for MNCs. Did that in South Africa and in Thailand and eventually landed in kind of like gravitated towards the data space and landed at Gojek in Singapore. So at the time, this was 2017, Gojek was about kind of a billion-dollar valuation company. It’s an Indonesian company. It’s a ride-hailing slash digital payments slash food delivery kind of super app. And so they have like a logistic network, kind of like Uber and Lyft, and they power various services through that logistic network. But so basically, we had a directive at the time to start a data team.
And the company had, I guess it was 4 or 5 core problems that they wanted to solve. And they knew they had a lot of— so they had engineering teams focusing on the product itself, and they knew they were sitting on a lot of data, and they wanted to improve the product experience for those core products using the data that they had been building up. So they staffed a data team, and I led the engineering for that data team. And they also hired a bunch of data scientists. I think in 2017, the kind of idea was you just hire a bunch of data scientists and you put them in a room and then you something’s going to happen. A lot of companies did that. Turns out it doesn’t really work that well. But yeah, so at the time, we had problems like, you know, we had pricing.
[02:53] Tim: One of the—
[02:53] Willem: it was one of our key problems, like how do you price a ride-hailing company? Like, so the supply and demand is obviously a big factor there. But if you want to make a booking from one side of the, you know, the city to another, what price do we give the user? Matchmaking, which drivers do we assign to which customers? We had things like fraud detection, that was a big one, and food recommendation systems. And so those are all kind of ML systems that we wanted to build and that we eventually did build at Gojek. So when I started there, we were building these solutions essentially. We were not a kind of platform or product team. And then as we built those, our team was about 5 to 10 folks. The data science team was growing rapidly. It was about 30, 40 folks.
And at some point, we realized that we could not maintain, you know, keep up with all the data scientists and all the use cases they had and the speed that they were iterating at. And so we kind of pivoted towards building reusable tooling and eventually platforms for them to operate on. And when we started looking at the platform, building platforms, we looked at the complete ML lifecycle and focused on the area that we thought was the lowest-hanging fruit. Like, what was the central service or central tools that we wanted to build along this ML lifecycle that could really address, you know, what was really slowing down our data scientists and their ability to deliver solutions for the company on those core products? And so the feature store is one thing that we built that was a key problem that our data scientists had.
That was like, how do you take data that you have developed, or how do you work with data in the offline sense, and then productionize that data for your ML use cases? So, there’s a lot to talk about there. But so, Feast was the product that we built and then eventually open-sourced. We built that in collaboration with Google. And we deployed that and operated that. It’s still running today. Gojek, and it addresses, I think, the majority of the ML use cases, although not all of them. So that’s— Feast is the feature store that we built with them. And yeah, so that’s kind of the whirlwind tour of how I got to where I was.
[05:11] Tim: And so we’ll dive into some of that a little more, but then you recently moved to Tecton, which is— are they supporting your work on Feast or are you working on the Tecton product as well?
[05:26] Willem: Oh, that’s a good question. Yeah, so I’ve just joined. I’ve left Gojek and I’ve joined Tecton. So Feast entered the Linux Foundation, and so it’s kind of like an open project now. It’s not governed by a single company. And my focus right now at Tecton is almost entirely Feast. So Tecton and, you know, the Feast developers and community really believe the same thing, like we can build the best-in-class feature store. And so I’m focused on Tecton, the core product, as well as Feast. But we believe that, you know, together these 2 solutions can over time address the same problem space. So they’re not— they’re separate products right now, but they’re kind of converging to one thing. But most of my attention—most of my attention right now is on Feast. And so they’re investing heavily into Feast. So injecting engineering resources and product resources and helping us grow the community, grow the user base. So that’s something I’m super excited about. But yeah, to answer your question, I’m working heavily on Feast.
[06:30] Tim: Feature stores are something that we’re hearing a lot about in the last few years in the broader data science, machine learning space. But I think a lot of people are still, me partially included, are still wondering What is a feature store and particularly how does it distinguish itself from other types of data sources like maybe just your company data warehouse or transactional databases? So what makes a feature store unique in what it provides?
[07:02] Willem: I think, well, it’s kind of like a Rorschach test. Everybody has their own idea of what a feature store is. Some people, if you give them a Git repo with a bunch of PySpark transformations, they say that’s a feature store and others say Redis cluster is a feature store. So, depending on if you speak to an engineer or some data scientist at a bank. My definition is, to me, a feature store is an opinionated operational data system specifically for machine learning. So, there are some unique aspects to what makes a feature store a feature store in my view. One is it connects the 2 worlds of offline and online. So, you’re development and production environments, as well as the kind of data and ML worlds.
So, on the one side, you have—one side, you have folks creating features, working with data, engineering data, and on the other side, you have consumption of that data. And often in machine learning, especially in the kind of MVP cases or kind of nascent cases, you’ll see teams working with, or not decoupling these 2 processes. They’ll build end-to-end ML solutions that essentially, you know, you’ve got one pipeline that transforms everything, trains the model, deploys the model. And what we found was that was extremely hard to iterate on for different teams. And it led to a lot of siloing of transformations and duplication of work. And so the feature store essentially decouples those 2 worlds.
So you can have one group creating data or features and the other group consuming them for training models or serving models. So that split is one of the key things and that, like, on-ramping, the ability to on-ramp your data into the production world is a key value prop. And then there are some other things like, feature stores provide a unified view between kind of the training and serving worlds. And I think that’s important as well, but I guess less so than the ability to kind of operationalize data. Yeah. So to kind of compare this to like a warehouse, a warehouse doesn’t have an online view, right? So you can’t query a Snowflake or something in production.
[09:25] Adam: When you say an online view, you mean something like an API or a service that’s like a low-latency thing that says, give me the value of this feature right now at this exact second.
[09:34] Tim: In reading the Feast documentation, it brought me back a lot to the Google paper, Machine Learning: The High-Interest Credit Card of Technical Debt, which I assume you’re familiar with. Is there— do you think this movement towards feature stores in the last few years was motivated, at least indirectly, through that recognition that you can create a real mess through jumbling all these kinds of pieces together?
[10:06] Willem: Yes, definitely. So what we found was that often you don’t have, you don’t know what the best design is. You just realize that if you look at the landscape of systems that you’d built, there’s inefficiency. You have an intuition for that, but you don’t know how to properly solve that. So we actually iterate on a lot of different architectures. And if you look at like TFX and a lot of the other DAG-building solutions. That’s a kind of different approach where you kind of— the transformation gets ported into different use cases and solutions. The feature store architecture is just one way that you can solve this problem. But it’s— so to go back to your point about that high-interest credit card, yeah, that definitely resonated with us. We saw that there was a lot of technical debt in our solutions.
This architecture is one of the ways you could address that for specific the data, kind of operationalizing data, which we saw as the highest debt, right? So, we just had, like, incredible amounts of duplication. And it’s also just, it’s so many things in one. It’s like, you know, you’ve got teams that they didn’t work in a structured way and just within their own team, or it could be just a single person, you know, duplicating his own work. They didn’t have a way to organize their data and to iterate on their own versions of their models or solutions, like end-to-end ML solutions.
[11:39] Adam: I can see feature stores showing a lot of value at the point where you get this duplication between teams. Maybe you’ve got the ride pricing model versus an ad placement model or whatever. And maybe these teams are duplicating data or duplicating code or duplicating whatever. But You just mentioned a single developer on a single team working on a single product. Do you see a time in the ML product lifecycle where it’s time to start investing in setting up a feature store like Feast?
[12:13] Willem: You mean for a single person?
[12:15] Adam: I mean, so let’s say you’ve got some hack week project and you’re like, hey, look, I built this little model to do something. You might not want to invest in a fully featured offline/online feature store to do that little product project. But then when you’re like, okay, let’s make sure that this thing is rock solid and going to all of our users, that’s when you kind of harden up some of your infrastructure and perhaps a feature store could be part of it then. Or perhaps you could stick with this sort of single end-to-end pipeline, like you were saying you had at Gojek. And at some point, a feature store should come in. And I’m curious where you see that point.
[12:52] Willem: Yeah, that’s an awesome question. So that’s very topical because This is exactly what we’re working on right now with Feast. So one of the— we spoke to a lot of our users, and it turns out that almost all of our users are platform teams, or, I mean, our adopters, or kind of like, there’s like 5 to 10 engineers working in a centralized data platform or ML platform, and they deploy Feast, and then their users kind of use that. So if you look at, if you consider that to be like a ladder, there are almost no lower rungs to that ladder. If you’re a single data scientist, I believe that there is value to having a structured way to get into production, or at least to organize your feature engineering so that when you do go to production, it’s just like flipping a switch or something.
But most feature stores are not organized like that, and they’re not easy to deploy. And Feast is no different. It’s kind of hard to deploy. It’s Terraform and Helm, and you need to have some Spark configuration or Beam, depending on which version you’re running. So typically, the folks we’ve spoken to, if you speak to platform teams, they’re happy to do that. And it’s a once-off cost and everybody benefits. And as a data scientist, you typically don’t use Feast at the start. But as soon as you think you’re gonna start operationalizing or productionizing your system, you start onboarding. But what we’re trying to do now with Feast is, really ask the question, can a single user find value in Feast?
Or at least, can a small pod, a solution-oriented pod, like, if you’re just solving a problem and you don’t want to build a platform, you don’t want to serve other people, but you want to lift a key metric or something in the company, can you deploy Feast? So, our focus right now is heavily on, like, how can we make it more lightweight? How can we make it easier to deploy? And it’s kind of an experiment right now, but That’s something that we do think is possible and something that we think will kind of expand the reach of Feast if we can pull it off.
[14:48] Tim: Can we dive into that a little bit more and talk concretely about if a single user was going to start dabbling in Feast, what— I guess I’m a little bit interested in the setup process, but more so interested in what functionality is actually going to be offered and how it fits into their actual machine learning pipeline?
[15:15] Willem: Yeah, so I think one of the key things to highlight about Feast is that it doesn’t provide you a means of doing transformations. So that’s something that I guess is a little bit controversial because most feature stores do provide that. So Spark or SQL or some kind of transformations, but Feast mostly allows you to do kind of the serving aspect. So It provides a unified API to kind of train your model and serve your model, and it provides a way to abstract the ingestion of data into the store. So whether that’s offline store, online store, used for training or serving. And then we have some validation capabilities there.
But to drill into what a single person would find valuable there is— so the idea would be that this person does want to productionize their kind of, like, let’s say they’ve got some end-to-end ML solution that they’re building. They want to productionize that. The use case that we’re thinking about is kind of a, let’s say, a small team. Let’s say it’s maybe it could be 1 person, it could be like 1 to 3 people. They’re working on some kind of basic model. One of the biggest problems typically that data scientists face is that They want to integrate with an engineering team and the engineering team doesn’t really want to help them. And the more the data scientists can do, the better and the further they can get to kind of productionizing.
And if they can just make it a single API that an engineering team can integrate with, that makes it a lot faster for them to kind of go get live. So, what I’ve often seen at larger companies or some companies is that there’s some kind of serverless environment in which a data scientists can deploy like a PyFunc or some basic model serving. And so Feast would be in this easy deployment mode. What we want to do is provide a way that Feast could slot into that environment without you having to deploy like a Kubernetes cluster or any kind of large-scale Spark environment. And so what we’re trying to do now is reuse only the infrastructure that data scientists already have available to them. So basically, instead of Feast being large-scale infra, kind of rework it into workflows that are executed from an SDK.
So this might be kind of, you know, it’s kind of a large shift in like how Feast is currently organized. But so basically, you’d use Feast out of an Airflow pipeline or existing ETL system, and it would orchestrate data movement let’s say, move data from your existing warehouse. Let’s say you’re a data scientist and you’ve got dbt, you’re doing feature engineering with dbt, you can then post that process, move your data into a, let’s say, into an online environment. That could be a Bigtable, a Datastore, it could even serve out of a bucket. And then you deploy, you know, your model serving and it would have a Feast client that reads out of that store.
And you’d be able to also ingest, let’s say you’re doing some ETLs in your data transformation systems like Airflow or whatever else you’re using data flow, you’d also be able to ingest into your offline store. And so Feast would facilitate the organization of that store, structure the tables and allow you to ingest the data and allow you to synchronize the data into your production environment And what we want to do eventually is also provide safeguards. So kind of metrics and statistics and validation, all those things that prevent you from sending the wrong data to your models in production. But I think a key thing there is we would not ask the user to deploy new infrastructure. And the only thing that they need to deploy is the online serving.
And we try and give them a kind of like a basic in-memory mode that doesn’t mandate kind of a big table or something of that scale.
[19:21] Tim: So in the context of Feast, then, what is a feature? What makes something a feature and how is it defined?
[19:29] Willem: So in the context of Feast, that’s pretty much any data point. So it’s a column essentially on some entity that you ingest into Feast. or that you materialize into Feast. So we don’t really differentiate on what creates it or between data and features because everything that’s being fed into Feast is essentially a feature.
[19:52] Tim: And so by saying Feast doesn’t do transformations, you mean basically those columns have to be created somewhere else if there are transformations of the original data? Yeah.
[20:04] Willem: So your transformations are happening directly upstream from Feast and that’s kind of A byproduct of the design that we had of our larger data infrastructure at Gojek. So we had stream-to-stream transformations and batch-to-batch BigQuery transformations upstream from Feast. And so Feast was only the layer that’s ML-specific that we’d use to operationalize that data.
[20:25] Adam: I’m curious, when you’re talking about these transformations, is this something like, let’s say the feature is the number of rides a person took in the last 3 days or something like that, which is an integer. And so Feast expects that integer number to be sort of inserted into the store.
[20:45] Willem: Yeah, that’s exactly it. Yeah. So it’s transformed raw event data in most of the cases. Yeah.
[20:52] Adam: And then there’s a post-transformation step that you could do where, say, the feature in your model is a one-hot encoded day of the week. in that situation, I would probably want to store either the string day of the week, say, or a timestamp or whatever. And then after fetching that back out at inference time in the real-time fetching, do the one-hot encoding with my own logic, whether that’s TFT or whatever.
[21:23] Willem: Yep, that’s right. So some of those transformations we’d leave up to the model or whatever the modeling the model unit is. It could also be the model serving, but it’s not handled right now in Feast. So on the Tecton side, interestingly, they have a specific thing called an on-demand or real-time transformation that allows you to define those and apply those for both training and serving. But we don’t have that on the Feast side yet. So those are basically transformations that happen that cannot be precomputed, right? That’s what you’re saying, right, Adam?
[21:55] Adam: Yes, I believe so. Yeah, I’m really interested in how the feature store handles the transformation things. One thing that I’ve seen in my career is that if you enable engineers or data scientists to do some sort of computation within an environment, like within the feature store, say, they’re going to do much more expensive things than you think they’re going to. They will immediately hit your guardrails. And so if you tell your customers, Hey, we have sub-20-millisecond fetching time, but then they want to do this transformation that takes 8 seconds every time they request the data. We don’t have to go into it now, but I’m just curious how that is a thing that is dealt with.
[22:44] Willem: So that’s interestingly not really a big problem for Feast. I mean, exactly how you said it, because we don’t do those transformations. They’re all upstream. So everything that’s being ingested into Feast is already computed. But we do have kind of variable length arrays that you can store in Feast. So you could store a string, for example, or a list of strings in a feature. So sometimes data scientists will just give us like 1,500 lines of JSON as a single value in a single, like, in an array of like a single cell of a feature. And so it’s extremely difficult to provide any kind of guarantees on those kind of features if you don’t have, like, fixed-length features or feature sets. And so, we’ve run into those problems before.
[23:38] Tim: Yeah.
[23:38] Willem: So, I think Feast, our SLO at Gojek was 10 milliseconds, and we ran that off of a Redis cluster. But, you know, sometimes you’d violate that if the data structures, the values you were storing, didn’t conform to those basic requirements.
[23:54] Adam: Yeah. I’m curious what you do in that situation. Do you try to work with the data scientists to say, hey, let’s figure out how valuable this thing is to you and if it’s worth the latency, or do you just sort of tell them it’s gonna be late? I’m always interested in that interaction between an infrastructure team or project that would be managing Feast and the internal customer team that would be consuming it.
[24:20] Willem: Yeah, I think this is the point where we should probably highlight that these operational data systems have multiple teams working on them, and they all have different incentives. And a lot of the problems you’re solving with these systems are organizational. So, data scientists often don’t care at all about latency. They just want to kind of ship their new data, and, you know, they want to see the uplift in BCR or whatever conversion rate that you’re looking at or whatever metric. But the engineers typically care a lot because they’re the ones integrating with the model serving and the model serving depends on the feature serving. And I think data scientists only care as much as like, what’s the ratio of feature retrieval and feature computation versus the actual inference, because then they can use slower and fatter models, right?
So at least that’s my experience. But our approach is basically hey, these are the dimensions that could slow down your feature retrieval. And we say, like, okay, if you ask for more entities, it’s going to be slower. And if you ask for more features, it’s going to be slower. If you ask for more elements in your array on a specific feature, it’ll be slower. And so we give them all of those and say, here is a baseline for you of what you can expect. So let’s say 10 features, 10 entities, single value per feature. and the type is this size, then you can expect like 5 milliseconds, maybe 99p is 10 milliseconds or something. And if you double everything, it’ll linearly slow down your retrieval. And so we gave them like a baseline and they could estimate their SLOs off of that. But it’s like a rule of thumb.
It wasn’t a hard guarantee, but it was good enough at the time. But I think it’d be much better if we could give them like a ceiling and then the engineers wouldn.t have to kind of load test a specific set of features before going live.
[26:08] Adam: I’m curious, I wonder if we should talk a little bit about the offline part of Feast and how that works. I’ve got the Feast website up and it’s got this function called get historical features, and you give it some names of the features which are registered in Feast’s feature registry. But if I say, give me these values of these features, can you talk a little bit about what happens behind the scene when I make that call? And ultimately, I get a DataFrame or something back, but what executes when I do that?
[26:39] Willem: This is a very tricky one to actually explain if you want to go into the point-in-time correctness and the time series aspects, but I can give you the high level. So it depends on which version of Feast you’re looking at. If you’re looking at the versions 0.1 to 0.7, that was based on GCP and It ran— I’ll talk about that flow because it’s easier to explain, and I’ll talk about how it’s changed recently. But essentially, we store all data in BigQuery tables, and then you’d have a single query that would— the user would provide an entity DataFrame or a spine DataFrame, and that contains— so basically, let’s say you’re doing a query on features for drivers. the user would load a DataFrame through this get historical features that contains driver IDs and timestamps.
And those are basically the timestamps that correlate to observations that their model’s being trained on, and they just want to enrich those with features. So this gets loaded into BigQuery, and within BigQuery, Feast already has organized all of these feature tables, and so we call those feature sets. So a feature set will just have many, many columns for a specific entity, and in this case it’ll be drivers. So maybe you’ll have like the driver’s profile feature set and the, the drivers, you know, for ratings or, you know, movement or, you know, whatever that feature set is. It’s just a grouping of data features that all occur on the same timestamps or the same events basically.
And what Feast will do is it’ll do a query on— so for each timestamp, so for each entity and timestamp pair basically, it’ll do a scan backwards for each feature that you’re selecting out of each of those feature sets, and it’ll join it onto the entity DataFrame or the spine DataFrame. So you’re doing basically like a join onto all of these feature sets, and you’re scanning backwards and ensuring that it’s kind of like a fuzzy time join. You’re finding the latest feature value for each one of those timestamps, basically those historical events, and enriching that with features up to a certain point. So you don’t scan back infinitely, you scan back up to the maximum age, otherwise it’ll be an extremely costly operation. And then you export that DataFrame and return it to the user. So that’s kind of the high-level process.
[29:14] Adam: I was just going to ask about how long that takes for a, you know, not for like a massive data project, but for like your sort of standard users.
[29:21] Willem: I think the query takes about, depending on the size of the data, I guess like 10 to 30 seconds.
[29:28] Adam: That’s great.
[29:29] Willem: Yeah, but just to be clear, it’s in BigQuery, and so you can kind of infinitely scale with BigQuery. Most queries take 10 to 30 milliseconds. If you ran that on Spark and you didn’t optimize things, it could probably take, like, hours.
[29:46] Adam: And so the process that you sort of— like, feature sharing is a super big topic kind of in the industry right now. And so it sounds like there’s these features about drivers’ profiles, things like that, that some one team or multiple teams may contribute to, to say, you know, this is the driver’s, I don’t know, kind of car and where they live and whatever other profiles file information, and that can be sort of shared generally. But if I’m trying to detect, you know, fraudulent drivers and you’re trying to detect how long it will take for them to get to a pickup or whatever, those spine datasets are the things that I need to provide for my problem and you need to provide for your problem. Those have to be put into Feast sort of first, and then we fetch everything back out and it does this join for you. Is that— Did I interpret that right?
[30:35] Willem: Yep, that’s right.
[30:36] Adam: I like that it sort of breaks apart the idea of what is universal within the company or the organization who’s sharing all of these things and what is specific to a model or a project.
[30:48] Tim: Yeah.
[30:49] Willem: So the features apply to different use cases and so the spines will be different depending on what you’re training for. Yeah, you’re right. So that’s kind of part of the motivation. The means in which we’re fulfilling that, I mean, allowing the user to provide a DataFrame could be different. So I’ve seen Zipline and some of the other folks will ask the users to provide a query and that’ll produce kind of the entity set or candidate generation or whatever you call it. And then you’ll kind of enrich that with features.
[31:19] Tim: If I can drop back to a very practical level again, say I’m a data scientist. and my company has Feast in place, and I just want to train a simple model in scikit-learn. What is that process going to look like on a— whatever level of detail is just above writing code? Like, what’s the high level of the code that I write then to access that data to train a model?
[31:50] Willem: Would you also be the producer of the data or just the consumer?
[31:54] Tim: Well, let’s say for now I’m just the consumer.
[31:56] Willem: Oh, okay. So basically, you’d explore what features exist within Feast. So you normally start with some business use case, right? And it’s either customer-centric or driver-centric or song if you’re Spotify or whatever entity you’re looking at. And you know what data you’re going to get and what features that can be enriched with. and what model you want to train. And so you’d look at features that exist on that entity. You’d probably export like a kitchen sink of those. So you’d look at that through the Feast APIs.
[32:31] Tim: There’s not a GUI, that’s all through code that you’re gonna—
[32:34] Willem: Yeah, that’s all just through the kind of SDK.
[32:37] Tim: Yeah.
[32:38] Willem: Yeah. So then you just, you’d export, you know, you’d select a bunch of those features. We don’t have a good process for filtering that down right now. I mean, we allow teams to annotate those features. And, you know, you could add labels and you could kind of group them and you can slice and dice that a little bit. But there’s no way to say these features will be great for your use case ahead of time. You can look at what other models have done that are similar and what features they’re consuming as a baseline. But, you know, to start, what often happens is that the data scientists produce their own data and then they will consume those and then they will enrich from other teams. But what we wanted to get to was that they would just start off by just selecting data on an entity, so features on an existing entity.
But you could do that and then you could export those features, you could train a model, you see which ones are the best. So you’d use this get historical features and you just train your scikit-learn model and then if you want to productionize that, it’s the exact same list of features and that same model that you ship into production. And then you’d call this Feast API again for those features in the online environment.
[33:48] Tim: And is that returning Pandas or NumPy, or how does that work?
[33:53] Willem: So Feast for training returns an Avro DataFrame, mostly because BigQuery produces an Avro DataFrame, but you can easily just— we provide helper methods to convert that into different formats. In production, We give you a dict for Python and for Java and Golang, it’s native arrays or lists.
[34:13] Tim: Java, Python, Go are all natively supported. Is that across the whole Feast stack? No, no.
[34:22] Willem: So Java and Go are purely for serving, so production environment.
[34:28] Tim: Yeah. So then the definition, the feature definitions and things are all in Python?
[34:35] Willem: Yeah, so everything is, you know, all our APIs are gRPC-based. And so the Java and Go are basically 2 implementations that we have implemented only in the serving environment on those gRPC APIs. But underneath, all the types are based on kind of those protos and gRPC. And Python, we’ve done the whole gamut, like everything from the management of features and ingestion and the retrieval.
[34:59] Adam: I’ve got one sort of last question about feature sharing and sort of either how Feast does it or your experience with this. But like, let’s say one of these features is a driver’s home address, which is, you know, highly sensitive information, and sharing that among a whole bunch of teams might be a little dangerous. Is like, do you have any sort of guidance on how to deal with this sort of sensitive information when sharing features, either Feast implementation specifically or just sort of in general?
[35:35] Willem: So, with Feast, we didn’t fully address that problem, but we provided some tools for those teams. So, basically, if you have an operationalizing layer like Feast on top of an existing, I guess, lake or warehouse or streams, often that layer is not the one that’s the source of truth for PII or sensitive data, right? So what we did at Gojek was we’d have a, you know, like a data warehouse team or a central data team that would annotate and identify that data and make sure that that’s secure. And Feast would allow you to isolate data within projects and then only expose that data. So basically within a project, that’s a concept that’s one level higher than these feature tables or feature sets, you’d not be able to access that data through Feast if you didn’t have access to that project.
So we gave you some crude isolation there, but Feast also, at least Feast 0.7, allows you to isolate kind of— you can have different serving layers, both for training and for online serving. off of the same kind of unified ingestion. So, a way to describe this, like, in the kind of whole universe of features that Feast has, you can have different subsets of those features consumed and stored within different environments. So, if you have, let’s say, in Gojek, we had like GoPay. So, let’s say payments or digital payments. And I think they even had a bank. At the time, just before I left, what we were looking at was giving them an isolated environment in which they could store their features. So, it’s a complete VPC.
It’s like a different VPC, but there’s still the same central registry of features that they share with the rest of the organization. But it’s just nobody has physical network access to that serving layer. So, they can’t actually access the data. They can just see the metadata of the data features. So we could restrict them on access time as well as kind of through the actual physical data plane. But what a lot of what we’re looking at right now is also implementing this for Feast, like, but what is the ideal solution and what’s the solution for Tecton as well? So we’re looking at ACLs and kind of surfacing or allowing the upstream systems to surface those PII, that information.
And the feature store shouldn’t really be the place where you define that as the source of truth, unless the feature is uniquely defined in the feature store and it only exists there. And it’s kind of rare for the feature to be the PII data, but not the underlying data. So, typically, that kind of propagates down.
[38:32] Adam: Right.
[38:32] Willem: So, you know, we’re, you know, like, if you speak to large banks or corporates, this is something that’s really, really important to them. But we’re trying to kind of delegate this to upstream systems and just integrate with those systems and try and not abstract and reimplement a lot of the ACL and PII and security functionality.
[38:51] Adam: Let’s see. I can totally see this be a very important consideration for a lot of companies and having that functionality. It’s great that it’s there or it’s supported at least.
[39:00] Tim: Willem, so say you’re a data scientist like me and you’re at a company that doesn’t have any type of feature store but feels the pain of the same type of problem that motivates needing a feature store. Do you have any advice on how we, these hypothetical people, go out and look at the options out there and consider whether something might be appropriate for our problems?
[39:36] Willem: Yeah, I think it’s, I always enjoy it when people come to us and say, like, they really know they need a feature store. The space is still a little bit nascent right now. And so sometimes people sometimes will try out feature stores, but they don’t really know if they need it. We see this a lot, actually. So, I actually, I mean, I’m happy when teams try and build an MVP solution without any kind of specific infrastructure like feature stores, and then they realize that there’s a problem there, and they try and solve it with a feature store. So, if you have, you know, if you’ve tried to deploy using Redis or some kind of other solution, and you, like, often it’s the teams that have, like, the real-time serving needs that run into this problem first.
Or if you’ve got multiple folks in the project or multiple teams on projects working on features, I’d encourage you to have a look at Feast. If you’re just a single person starting out, it’s probably not going to be something you need right away. If you’re just looking at, like, batch use cases, not, you don’t have online serving, it’s probably not going to be a high ROI for you either. So if you have just batch training and batch scoring, but if you have online, maybe try out getting your system up and running first. And if you run into kind of like reuse or scalability or kind of deployment pains, like, you know, data scientists can’t get into production and can’t get new features into production without engineers, then I’d encourage you to have a look at Feast. There are some other projects out there as well.
I’m very biased, so I’m not going to list them out, but yeah, you can look at featurestore.org, I guess.
[41:22] Tim: And you’ve talked a lot about the roadmap for Feast here, and we don’t have to necessarily recap that, but any things that are coming down the pipeline that you’re particularly excited about that you’d like to share?
[41:34] Willem: Yeah, I think the 2 things— so we’ve spent a lot of time thinking about what exactly will the modern data stack look like in 2 or 3 years? And it’s moving it towards this kind of like ELT world where, you know, a lot of managed services and the feature store seems to have some overlap with some of those upstream systems, like, you know, the dbts of the world, if you’re doing transformations or the, you know, the Sparks of the world. And so for Feast, we’re very focused on this kind of unification of the training and serving and the operationalizing data from those existing modern data stack tools.
So I think the things that excite me the most about Feast going forward is this kind of simplification of its architecture, slimming it down, having it occupy a smaller amount, kind of not just less infrastructure, but less responsibility and scope. But I think our biggest focus will be kind of making it easy to deploy. And I think I’m excited for folks to try that out. That’ll probably not land soon, maybe the end of, you know, I guess May or so. And we’ll also be investing heavily in the data kind of validation statistics, both for kind of batch and for the online production environments. So that’s something that I think I’m super excited to double down on.
We’re going to try and integrate with Great Expectations or TFDV or many of these tools, try and just re-leverage or try and leverage those existing tools to make sure that data scientists aren’t feeding garbage into their models in production. Which in our surveys with data scientists were almost always like the second most important thing just behind how can I do data transformations. So, I’m not sure if that also resonates with your experience, but these sort of data scientists in Gojek, they depend on other teams for their raw or intermediate data. And it just, it was the number one cause of failures for them was upstream data breaking.
[43:44] Adam: That definitely resonates with me. Yes.
[43:46] Willem: So, basically those two is the kind of simplification of the deployment and kind of the validation functionality are the two big things.
[43:53] Tim: Yeah, awesome. And obviously Feast has corporate backing, but it is an open source project. Can you share about ways that if people are interested in getting involved with the Feast core project, things they might consider?
[44:08] Willem: Yeah, we have a mailing list. We have— if you just go to feast.dev, you can see our kind of community and getting started pages. We’ve got some info on it there. We’ve got a community call every 2 weeks. And we also encourage contributions. We love it when people create issues, even if they say there are bugs or problems, or they just want to request new features. So just get us on GitHub or join our community calls or just jump on our mailing list and you can just stay up to date. So it’s feast.dev.
[44:41] Tim: Yeah. Excellent. Any parting words that you’d have? Nope.
[44:45] Willem: Thanks for having me on the show.
[44:46] Tim: Yes, sir. Thank you very much. And Adam, thank you for bringing your perspective on this as well.
[44:51] Adam: Yeah, thanks for having me.
