Folks, in your launch posts remember to describe what you're launching and what it's for. Ideally in the opening paragraph.
- Intro: Explains problem but not what is Neki or what it's for.
- Why Neki: Explains alternatives and why they suck but not what is Neki or what it's for.
- How does Neki work: Describes the technical components but not what is Neki or what it's for.
- What you get beyond sharding: Explains how to run/deploy but not what is Neki or what it's for.
- What is a platform preview + Try Neki today: 140 words on the definition of "preview" and links to get started... but not what is Neki or what it's for.
Edit: The landing page has 100x more useful information up front -> https://neki.dev/
Edit #2: They've since added a "What is Neki" section. Good reaction time.
I scrolled through the landing page and even the "Get started" page and I have still no idea what is this. Is this a managed db? Something I can run myself? No idea.
The landing page which is linked at the start of the second sentence of this page…? I don’t completely disagree with you, but this is mostly hypertext working as intended, is it not?
Yeah this meta commentary on HN is very confusing for me. The structure of the article is literally
- short paragraph saying "we launched neki"
- short paragraph saying why neki was needed
- paragraphs about what neki is (with a header)
That seems like a very reasonable structure for a blog post like this one. Not being able to get to the third paragraph of a blog post seems like a "you" problem, not a blog problem.
Why "sadly"? The person writing a link should provide enough detail for readers to figure out whether they should follow the link. I rolled my eyes the first time I saw this entry on HN because it just said "Neki" and the (generic) target domain name. A two-word description got added, and that makes it interesting instead.
It's sad because I think links are the best thing about the web, and it's increasingly clear to me that many internet users either misunderstand or actively avoid them.
Why do people post screenshots of an article on social media instead of linking to it?
Sometimes ignorance, but often because the web is such a hostile experience these days - paywalls and subscribe banners and cookie warnings and slow pages.
The top concern I've gotten from dev teams when proposing HA distributed postgres (e.g through RDS Aurora global) is that eventual consistency is not suitable for many workloads.
Does Neki solve for this, and if so how? My understanding of CAP theorem is that this basically requires some compromises around availability, but I'm curious as to what that looks like in practice here.
Assuming it works the same way as their MySQL product, Vitess:
The usual way to run it is that you partition your db based on something like a user, so that single user gets a consistent DB, but anything cross-shard may not be.
I know when I worked at Block, Cashapp was using Vitess and getting cross-shard DB writes down and functioning correctly was one of the major blockers to adoption. (though I just did tls management for vitess and didn't write any workloads on top of it, so my impression might be a bit off)
The consistency guarentee assuming repeatable read or higher on each shard is parallel snapshot isolation, which is technically still higher than postgres read committed.
If you have serializable on each shard and you chop transactions so you have no cross shard transactions, the application is technically still serializable, it just behaves like two independent deployments of the application for different shards
Same here, first thing I need to know when considering a distributed system is how consistency is handled. If it's eventual consistency, what is the replication lag like? If it's strong consistency, can their network handle that? What happens when a node goes down?
I love the offer which is running on the "PlanetScale Metal".
They claim to provide unlimited IOPS (I/O operations Per Second).
I would like to get some of these drives myself. Sounds like a technical miracle.
On the other hand claiming such miracles makes me wonder which other aspects in this are are actually a working miracle or rather a mythos/false marketing claim.
I love how the Planetscale CEO has been shitting on multigres[0], Supabase's equivalent to Neki, and gloating about how superior Neki is, and yet Neki is not even open source!
Sam, it feels like you're just playing games. It's hard to take anything you say or do seriously. Are you going to OSS this thing? Or have you decided to backpedal?
Congrats on the checks notes closed launch, I guess?
Sam is always shitting on others. For years he has been screaming how bad PostgreSQL is compared to MySQL. Until getting silent about this topic for a few months and then launching PostgreSQL support. Seems he had to realize than the market for PostgreSQL is too big to ignore.
No it isn’t normal, its just this one person. I don’t personally find this language particularly appealing, especially for a serious discussion on technology. But hey, its a free country, they’re a startup and its one way of grabbing the scarcest of resources: attention.
They both are characteristics of software that people consider when deciding whether to use it.
Open source being by far the more important of the two, since it means the user is more insulated by terrible corporate decisions thus making it more suitable for long term planning.
who said open source has no value? I worked at GitHub for 6 years i clearly love open source.
We maintain Vitess using our company funds. Why are you going after a small company for defending itself against companies like Amazon but not criticizing the reason we have to do this.
I would love to open source Neki and I have not said we won't.
People think comments like this really bother me as if I believe that I've missed out on any dollars at all from you, as if you have the buying power or ability to drive any workloads that would in any way move the needle on my revenue.
I want to let you know that I have customers that pay PlanetScale millions of dollars and have told me they came to the product because of my Twitter account. So you will never make me feel bad about this.
I love to see this but as a heavy Vitess/MySQL user I do fear the split focus from PlanetScale. Hoping to see continued improvements on the Vitess side as well.
Selfish doubts aside, congrats to Planetscale on the launch!
don't worry, we still give a lot of love to vitess. at the end of the day they are both databases. the beauty of doing both is we can take learnings from each product and apply it to the other.
I am on a 12 Muni right now that is plastered with a Neki ad, and I looked up neki.dev literally 10 mins ago to see what this product is about. Did they time the release of this bus ad with the launch date?
How are cross-shard joins and transactions handled? Are there query patterns that would require application changes despite using the same drivers and ORM?
cross shard joins break down into single shard selects, and are joined in the routers. There are some query patterns that are currently unsupported, but we are working on adding support for them. The biggest thing to keep in mind when running on a sharded database of any kind is decreasing the number of cross shard queries.
> This adds some undesirable side effects, however. First, the coordinator becomes the bottleneck.
Apart from being able to use any node as a coordinator (and you can load balance them to avoid having "multiple connection strings), there's a new pattern which effectively allows you to have as many coordinators as you want. They are effectively "data-less" nodes. We have devised and implemented this pattern in StackGres [1].
> Adding a shard with more resources for a noisy tenant, or many small shards for a wide shard space requires substantial manual configuration of not only the database servers themselves, but wiring them up together with Citus.
Adding nodes (infrastructure) is what operators solve. In StackGres, adding new nodes means editing one/two characters from your YAML file: the integer number that represents the number of workers that you have.
Shard rebalancing is fully built-into Citus as a UDF, which you can call (manually or in an automated manner) over Postgres protocol.
> Managing backups is also external to Citus, so operators still need to build the proper infrastructure
Agreed, but it's also solved (see distributed backups in StackGres [2]).
> PgDog is a spiritual successor to PgCat, both of which improve on Citus's architecture substantially.
Unsubstantiated why. I assume it's because of the assumption that a proxy model is superior than Citus. To which I have to say that Citus model is also a proxy model, where the proxy just happens to be Postgres, which unsurprisingly, speaks Postgres protocol. Sure, there are nuances that we could debate in this area and we can say that Citus is not a "pure proxy", but that doesn't lead to concluding that a proxy model is better --it's arguably not.
The first thing I want to know when looking at a distributed system is how consistency is handled, ctrl+f for "consis" on both the blog post and neki.dev has zero hits. So I can only assume it will have eventual consistency with a considerable replica lag for real-time systems.
I see no mention of foreign keys, or any other constraints, across shards. If, as I suspect, they're not implemented, it would still be useful but at the level of Spanner 10 years ago.
Are you one of the developers ? If so, you need to add a specific page explaining in detail the guarantees that this gives.
For example: point 08 says "Assign different tables or workloads to different shard groups" and point 02 says "Split hot shards as workloads grow". How do those interact ? Can a single table be split across multiple shards ? If so don't you need 2pc to enforce primary key constraints ?
The largest database to ever run on Aurora MySQL runs on PlanetScale, and the largest database to ever run on Aurora limitless also runs on PlanetScale.
We went as far as we could on GCP before PlanetScale. Spanner is amazing but also amazingly expensive. And CloudSQL, also amazing as long as you don’t need write scaling. PlanetScale (Vitess+MySQL plus their branching / deployment and monitoring tools) is just not a combination offered on GCP. And while I didn’t use Aurora, from talking to and exploring AWS it doesn’t really have this combination either.
Yet. But it’s built by the guy who created Vitess and cofounded PlanetScale, so while it doesn’t have sharding implemented at the moment I think it would be foolish to discount it. Plus it’s open source.
The PlanetScale team spent years improving Vitess with lessons learned from many different high scale customer workloads. We continue to invest in and improve Vitess. It's a substantial and growing part of our business.
Vitess today is not the Vitess of 2021. About 70% of the current codebase was written in the last five years, based on lessons we learned running it at scale.
Neki is built for Postgres from the ground up, by the engineers who did that work.
On the Multigres comparison, I'll admit we get a bit salty, mostly because "Vitess for Postgres" invites the comparison. But to date there's no public evidence of a working sharding implementation in Multigres. Their sharding design doc landed this week, and it lists cross-shard query planning and resharding execution as explicitly out of scope.
Neki does sharding today, including online resharding that switches traffic without a maintenance window. It isn't open source, but you can spin up a cluster and try it yourself.
That is what the “regular” PlanetScale is all about. It uses the Vitess open source project to do write scaling through shading. It is very powerful and customizable and stable. It was built to run YouTube. Former happy user here.
Many of the engineers who built Neki also worked on Vitess for years. But it's a ground up new implementation designed to specifically target Postgres.
Sam is not an asshole. His twitter is great, he's definitely opinionated, and if you haven't had the luck of meeting him you might misread his tone as grumpy. It's not. It's sarcastic/dry humour and honest.
Perhaps he's quite nice over zoom, but since very few of us are going to meet him in that context, the rest of us are left to judge him by what he posts. And as they say, you never get a second chance to make a first impression.
Based on what I've read so far, with his shit-talking the competition, he seems like... a bit of an asshole.
> he just comes off as funny on twitter once you've met him.
Never heard of the guy before today, but his dickish, edge-lord aura (on X and in this thread) makes me not want to meet him. If I had to choose between buying from him, or some other random dude, I'll hear out the random dude first.
Oh no, how am I ever going to avoid giving my custom to someone who has allegedly cornered the market on - checks notes - a db proxy? Give me a break.
I can think of at least 4 other ways of scaling horizontally without switching DB engines. I'd have to be mightily constrained to have this particular product as the only option.
Sure, I don't doubt your ability to do so yourself. I just question your ability to trust some random unknown figure with your data if you have enough scale. That's quite different, actually.
There's obviously more to the story than outsiders are aware, but personally it just leaves a bit of sour taste in my mouth when the tech and blog posts you guys put out are amazing.
Did anyone say you said anything untruthful? No, they just called you an asshole. You can both be an asshole and be right.
PS - Just some advice that I learned in sales 101 early in my career...focusing on shit talking the competition generally isn't a strategy that works well.
In case you were unaware, if someone's defense against being called an asshole is to say "well nothing I said was false" they're almost certainly an asshole.
- Intro: Explains problem but not what is Neki or what it's for.
- Why Neki: Explains alternatives and why they suck but not what is Neki or what it's for.
- How does Neki work: Describes the technical components but not what is Neki or what it's for.
- What you get beyond sharding: Explains how to run/deploy but not what is Neki or what it's for.
- What is a platform preview + Try Neki today: 140 words on the definition of "preview" and links to get started... but not what is Neki or what it's for.
Edit: The landing page has 100x more useful information up front -> https://neki.dev/
Edit #2: They've since added a "What is Neki" section. Good reaction time.
reply