Hello SaaS developers, I have a question about SLA...
# general
a
Hello SaaS developers, I have a question about SLAs. Say you are offering a storage service for structured data. So it's more like a KV store than S3. IN general, I've seen two types of availability SLAs for such services: 1. Aurora, SycllaDB, etc., define uptime as whether you can connect to the service at all over a minute window. The service is unavailable for a minute if you couldn't connect for the entire minute. As a customer, it's your responsibility to provide the logs, etc., to prove the outage. They have to verify it, and if they confirm the findings, they issue a credit. 2. Companies like Confluent, RedPanda, etc., while not quite offering queryable storage services, measure their SLA through an independent probe on an internal topic/table. If all reads/writes from the probe failed in a minute window, the service is considered unavailable for the entire window. The customer doesn't have to prove anything. Is there any type of consensus on which is the better model from a service provider and customer point of view. Is there a third option at this granular level?
Personally, I find it hard to pick between them. The client side self reporting option actually represents the customer experience, but you are not directly measuring it and putting the onus on them to report errors is not ideal. The internal probe is better from the point of view that you can alert yourself of downtime that counts toward SLA. However, it is not necessarily representative of the customer experience. All things being equal, I prefer the interal probe, but am curious on what everyone's opinion is
m
From a customer pov I want a query slas. As a backup #2 is my personal preference. Not just for not having to do anything. But also it gives me confidence that the provider is monitoring
I don’t like 1 because connection metrics don’t mean anything. The tcp port could very well be open and I could connect…but so what if data ain’t moving
a
Agreed. Do you know any databases who provide query SLAs?
c
Could you make the case that Confluent kind of provides query SLA's if you count a "produce request" as a query? And yes, as a customer I do want query SLA's, because technically with the Confluent SLA it's that all produce requests in a given minute need to fail for the SLA to kick in. But if there's 2 out of 6 brokers that are completely stuck, it's possible that with RF=3, min ISR = 2, acks=all some of my partitions could be very offline. But that standard is probably too hard for a vendor to meet (and I'm about to become a vendor, so I totally get it...degraded EBS volumes keep me up at night)
a
You could make the case, if you really squint. The reality is that queries and produce requests come in all shapes and sizes. Just because the request from the probe works, it doesn’t mean your requests will work. This is even more true for more general purpose databases. I still haven’t found any database vendor who provides generic query SLAs. In general such an SLA will be hard to define since queries can fall across a large latency spectrum.
g
I’m wondering if you could close some of the gap between the options by allowing customers to install the probe in their account. This way the probe is more representative of their experience and the burden of collecting the “proof of availability” is shared.
c
Our cloud service is still in the design phase, but we are leaning strongly towards having probes run in the control plane so that it has to go through the load balancer + K8s ingress + authentication + the whole shebang just like the customers' requests. If a customer is using a cloud service, they generally expect it to "Just Work." To me, having to run a probe yourself and then provide logs as proof of an SLA not being met is not "Just Working" (i.e. I've rarely heard of people run a probe to monitor S3).
👍 3
Obviously, that model where we make requests from the control plane to the data plane clusters won't work in certain "LH Dedicated" deployments where the customer requires VPC Peering. But LH Dedicated is a long way away, and those customers will be big and sophisticated enough that we can find a solution.
👍 1
g
Yeah. I think you are in the right direction.
c
That architecture was inspired by this blog post about the Strimzi Canary. We stole the name...actually the first project that my first hire did was to build "LittleHorse Canary". The entire Strimzi Blog is A+++ content, this is one of my top 10. https://strimzi.io/blog/2021/11/09/canary/
g
FWIW, I did have a probe monitoring EBS (and I know Netflix had one for EC2 as well). Exactly because of how AWS handles SLAs
a
What did the probe for EBS do?
c
Interesting...I think Justine Olshan did a talk about how Kora monitors + recovers from slow disk volumes at the Austin Current22? There was a lot of other stuff in that talk, and it may have been someone else...but I do remember that talk and think it was Justine
g
Open a file and write one byte.
👍 1
🤯 1
c
I have lost sleep thinking about slow EBS volumes because of that talk.
g
Amazing how much you can discover by doing something so simple on many disks over time…
👀 1
c
Found it! Slide 13 https://www.confluent.io/events/current-2022/whats-up-with-availability-in-kafka/ But this slide mentions monitoring the threads, not monitoring the volumes themselves, so it may be a different thing...
🙏 1
g
Really nice. She worked on the replication/ controller mechanism, so she described the work of her team (detecting issues from internal Kafka state). My team and the storage team had more external monitors to augment (and so we can open support cases with cloud providers and provide proper evidence).
🙌 2
c
Wow, almost a year later I finally realized the pun in the title of the session:
What's up with availability in kafka?
up -> Uptime 😂
g
Hahaha
I forgot to mention an important resource! The SLO book by Alex Hidalgo is useful and practical. Also, if you talk to him, he'll: 1. Send you an SLO card game. It is a bit silly, but also has great examples of SLO monitors and the problems they solve. 2. Spend hours on the phone with you helping you with specific questions and concerns about your specific monitoring needs. 3. Sign you up for Nobl9. Whether you end up using it is optional, but I was impressed with how quickly they respond to product feedback.
❤️ 1
👀 2
d
Sorry for resurrecting this thread, I'm just coming back from Holidays. Great callout on the Strimzi Canary post - a similar one inspired the name of the first iteration of the Cloudera public cloud API canary I made with a few folks - we called it FEF0, as FFEF00 is canary yellow in RGB 😂 Corporate made us rename it Cloudera-Deploy when it was open sourced. I'm onboard with the both answers here, though I would observe that we get great customer feedback at Tinybird from owning up quickly with option 2 by providing an RCA and how we're going to fix it going forward. It's a maturity thing though, and depends on the service you're offering, but in a Serverless world you need a good framework for easily adding monitoring for customer behaviors that might lead to a problem, or test coverage when upgrades or config changes might cause the system to misbehave - an immature company will think 'phew, we got away with nobody noticing that one' but the more serious company is always worrying that nobody has noticed yet and enterprise customers will eventually check their own logs if you don't monitor it for em.
g
Big +1 on "Providing RCA and how we'll fix things going forward". I love it as a customer. And I am a big fan of doing this for my customers. One thing I found important and non-trivial is to align the company around the right level of transparency and technical details in these RCAs. Relatively easy in a startup, but I'm always so impressed that large companies manage to do this well.
m
+1 I love a good public RCA. and I trust companies that provide them far more than the generic "No longer broken. We will fix the root cause"
💯 1
g
Private RCAs are important to get right too. You'd think it is a no brainer, but I can count on one hand how many companies do it right.
💯 1