The new `m7g` instance types look really good - a ...
# general
g
The new
m7g
instance types look really good - a bit cheaper than
m6g
and 3x the network.
💥 1
m
looks interesting, so few people understand the netowrking throughput is critical for a bunch of very common use-cases, they see CPU and Mem under-utilized and "I don't understand why it's working so slow, it's not even using X% of the CPU..." When your app sends/reads a LOT of data - network can become a bottleneck very fast, 3X throughput is significant!
🧠 1
💯 2
r
I’ve been burned by this. pre-committed into a lot of compute resources, but then got stuck waiting for pods docker images download for like 90 seconds each, and god help me if I decided to drain the node and reschedule. and that’s just basic nuisance, not to mention data replication
👍 1
m
You can't help the pre-commit, Cloud or On-Prem, it is hard to predict future hardware development or needs. The only "medicine" is to spend more initially (at least in the Cloud), and use that to learn your needs better, and still you can't predict the future cadence of instance types, it's a market bet. More than that, it's exactly the spread cloud and hardware vendors live on - can you even imagine the pricing negotiations between say AWS & Intel? Two giants who know so much, and have to decide on a price that works for both... Its the real reason for Graviton and all the other innovations - cut out the profit margin of the vendors from my own pricing... added value - vertical integration (like Apple), but only an added value, I think they were happy enough with Intel to begin with
r
yeah, i agree. just saying that i lost the cloud pre-commit bet, but pretty happy with my on-prem bet 🙈 wouldn’t have made better judgements without seeing what works and what’s not important from the cloud. (not to say it is the same thing as maintenance will be brutal)
i was amazed at how apple excecuted their arm transition. supposedly, they didn’t want to build their own cpu and asked intel make them a better cpu, but priorities and margins, heh.
💯 1
c
Thanks for the pointer Gwen! We will look at using
m7g
for the Kafka brokers, and
m6a
for the LittleHorse Servers. The LH Servers are actually bottlenecked by CPU, not network; Kafka tends to be bottlenecked by either storage or network. I love K8s. Makes it so easy to do stuff like this 🙂
g
Not too hard with Pulumi + ECS either 🙂
👀 1
c
How hard is it to run Kafka on ECS? Strimzi makes it super easy to do so on EKS, especially with the Topic Operator and User Operator.
g
I haven't tried tbh
But I'm running my infra on ECS and swapping machines has been fine
✅ 1
s
Sounds pretty good to optimise your instances for the resource you make the most money out of :)
m
@Gwen Shapira you made me curious, why ECS? It's nice, but it doesn't really compare to K8S.
c
I agree about:
it doesn't really compare to K8S.
One other consideration though is that EKS doesn't really compare to GKE, OpenShift, Tanzu, or even some self-managed K8s offerings. My biggest complaints: • A Kubelet on a node can crash, but the EC2 instance is still reported as RUNNING and it never gets replaced in the group. • When an EC2 instance crashes, sometimes EBS volumes get stuck to it • Gwen mentioned EKS was a third-class citizen with regards to gp3 support
m
Yeah, but AWS had EKS which is managed and priced about the same (not sure exactly).
👍 1
g
"priced about the same"? Not sure same as what exactly...
m
Gp3? Isn't that EBS volumes? How is that related to EKS?
ECS vs. EKS
Yay, restarted discussions 😃
g
ECS doesn't make me pay for the 3 "master" nodes. So at my scale, it is far from comparable.
👍 1
c
Ah, I understand the confusion. I was suggesting that perhaps ECS is more attractive than K8s for an AWS user because ECS is a far more mature offering compared to EKS. And you're correct about GP3, it is a storage-related thing. Mounting Persistent Volumes on GP3 is hard.
Doesn't make me pay for 3 master nodes
I thought it was only $75/month for the control plane regardless of cluster size, because the control plane is actually multi-tenant behind the scenes?
g
In general, ECS APIs are much simpler to work with. I'll probably run into limitations sooner or later, or I may need multi-cloud. But so far, I'm good.
c
Yep, $876 per year for EKS control plane.
g
I can double check EKS pricing, but last time I tried, I ended up paying more than 4x what I pay on ECS.
😱 1
c
The kubelet + kube-dns + kube-proxy is additional overhead on each node, perhaps ECS is more efficient. However, the kubelet is pretty small...
m
That makes sense, AWS indeed give their in-house technology at a slightly better discount - and also, for a simple setup it works great, I started with ECS+Elastic Beanstalk and was very happy for more than 2 years. Then we switched to K8S and never looked back, but the scale was significant
ECS has a bunch of agents itself, but yeah, I don't think they work as hard as the K8S apps
g
Does AWS have blueprints for EKS? Having pre-existing scripts with all the subnet and routing figured out was a big deal for me too.
m
K8A provides pretty good helm charts, that do it all. However, K8S permissions model mapping into AWS (for accessing the cluster ) is apparently a headache
g
I need to re-check Helm. I heard it got a lot better after the transition to v3 and the server-based model. But its been a long time since I actually had to use it.
c
https://registry.terraform.io/modules/terraform-aws-modules/eks/aws/latest This is preetty good. We're prototyping with it (And from our previous conversations, I know you use Pulumi, which is compatible with such modules 🙂 )
👍 1
Anyways, it seems like you're quite successful with ECS. There is no reason to use K8s unless you have a reason to use K8s. Migrations for the sake of migrations are not great for the bottom line (that was a vacuously true statement, but I think you get the point).
r
@Gwen Shapira are you running your own k8s masters and rolling VM worker nodes manually?
does it allow CSI? i thought they only provide CSI under managed K8S
g
I believe I did. But I was not aware of an EKS alternative. I literally followed whatever instructions I found when googling for "EKS getting started"...
m
I'm not selling K8S here, I totally am for saving and least effort for max results...
g
Oh, I do appreciate all the suggestions. I know we'll be on K8s sooner or later. Just a matter of time.
c
I'm not selling K8S here
Hah! K8s and Kafka are my two favorite things (along with cold brew, Spyglass Golf Course, and my childhood piano), and a few weeks ago I saw a medium post "K8s and Kafka can get you fired". I took it a little personally 😂
😂 2
m
Regarding the TF Module - the difference with helm charts is that they usually come with more same defaults... Or rather - opinionated While the module needs to support every crazy use-case that you might want
c
Wait a minute. How can a helm chart be used to bootstrap a K8s cluster? I thought helm was to install something into a k8s cluster that was already running?
m
@Colt McNealy Wow, really? I think it's fake, never saw anyone getting fired for one of those, the opposite is true - more people got hired to manage it
r
@Colt McNealy I think it’s wrong to consider kubelet + kube-dns + kube-proxy as overhead. it’s a small price for what K8S provides. and the solution to fixing the overhead price is just using large instances. I never understood using 2cpu/4gb ram instances anyway. for large instances, the overhead is negligible
➕ 2
‼️ 2
m
@Colt McNealy Helm charts do everything - but they essentially have modules underneath Like other operators on K8S - the sky is the limit as to what you can make it do - the operator provisions stuff out of the cluster directly in AWS Crossplane is taking it a step further and makes the cluster the reconciliation engine for everything
g
@Rauan Mayemir if you don't have a lot of customers, large instances increase your baseline per-region cost quite a bit (as does the "k8s overhead"). And for dedicated clusters, this overhead sets a minimum cost of a dedicated cluster. Considering that some of us compete with AWS and their "your margin is our opportunity" attitude, this is really tough. Not impossible, but challenging.
r
I see. use cases surely vary
👍 2
m
Check out AWS K8S operator https://aws.amazon.com/blogs/containers/aws-controllers-for-kubernetes-ack/ I'm almost certain it works with Boto underneath (no TF state)
c
Ah, I think I get it. So there has to be a K8s cluster running somewhere, and you can use a helm package pointed at that bootstrap cluster (cross plane?) to create a K8s cluster somewhere else in the world?
👍 1
g
interesting
c
Oh my goodness if that can save me from using terraform or pulumi ❤️
m
That's one option - or you can install you cluster with crossplane with some other automation (how many times do you create a new cluster?) And manage the account with it. We are probably going for the crossplane per cluster and not a shared central one because that causes concerns of SPOF and possibly delays (it can be busy with some other cluster and not be as fast as you want)
You'll probably love Pulumi once you start working with it, but the learnimg curve exists It does have the advantage of being CODE which you can write in most of the popular languages today (java, js, ts, python and more)
@Gwen Shapira @Rauan Mayemir If your env is small enough even ECR is an overkill, just use EC2 auto-scaling groups, an even simpler API and very easy to reason about... A cluster is only really useful for two things imo: 1. The apps start faster because it's just a pod and not a whole instance that could take minutes to ramp up 2. For cases where you have multiple apps running (services) and then you want better utilization of resources which points to large(r) machines - this of course depends on your specific case. But there is a difference between medium, large, x large and 16xlarge (not to mention network throughput which grows with the machine size, and is maybe the only case where pricing isn't "linear" - CPU & Mem are on a exact X $ times - so 2 X small equals medium in price and power but networking not...
r
@Moshe Eshel I had a funny realisation that when all I do is code, the last thing I want to do is write more code for infra. and by code I mean anything harder than yaml or at least dsl (e.g terraform). so I stopped using pulumi for new stuff and transitioning to terraform for critical infra, and then gitops for all the other stack. (fluxcd in my case, but argo is just as fine)
❤️ 1
I wrote pulumi code in typescript and it was nice. but for some dumb reason I tought that rewriting it in Golang would make it super fast, so I ended up with an extra cycle of very slow build (due to golang’s lack of generics and insane amount of generated boilerplate, fixed long ago). now I just want to copy paste yamls 😂
👍 1
m
@Rauan Mayemir I can sympathize ❤️ I don't expect product teams to write Pulumi code, moreover this can actually cause many problems of multiple different ways to do the same thing. I think an infra team writing the code with support of operations (DBA, Network, SecOps) to cover the edge cases - and write a thin declarative layer on top allowing teams to "define" which resources they need, and the shared pipelines will execute the cod
r
but this is purely anecdotal. i think no matter what you use, at some point you’ll get jaded anyway
👍 1
m
Regarding making it fast, it can be slow as hell, i don't really care , it's not my product that is working slow, it's just provisioning
r
oh, that is so true. ultimately, i didn’t want pulumi because i’m not a devops guy. but when you scratch off the fancy devops tag, all the hardcore sysadmins use ansible and probably not very keen to start coding in a way we backenders expect to see code.
👍 1
💯 1
m
Yes, exactly they like Ansible and TF and maybe terragrunt they don't want to change... And your observations are so true, I sort of hate writing code, and am totally jaded regarding "new" technology (in our space where everything is some rehash of a 1970's paper), but, I love solving problems, and helping others solve them - and that keeps me going (and paid!), So I'm actually happy to go into work most days, and super charged when a solution I helped bring to life actually works and is being used.... Anecdotally, I have a friend that leaves her job about once a year, sometimes a bit more or less, because she gets bored, and is looking for a new challenge, but all she keeps on doing is coding in software companies (usually a SaaS of some sort), and I keep telling her, that's not the way
💯 1