I do a lot with cloud providers for my customer's products and have worked with Digital Ocean's products once before. I didn't have a particular opinion on them[0], and there's some things that seem off about the twitter thread when placed against the incident update report that this thread is linked to. So, all of that to say, I'm giving Digital Ocean the benefit of the doubt.
There are, however, a few things that could be improved about this process:
> Peer review of account terminations. For any account appealing a lock, two agents will be required to review the submission prior to issuing a final deny.
The devil is in the details. Do this in a manner that the person confirming that the account is committing fraud is unaware that they are confirming another's denial; otherwise "dude, can you approve that termination I just did? I want it out of my queue/that guy was a dick." is a risk.
Off the top of my head, I'd probably generate two support tickets (linked, but without that link presented to the account termination CSR team member), assigned directly to this person, hidden from others ("hiding", along with training/process improvements is likely enough). If one person disagrees with the termination, close out the other person's ticket. If the CSR sub-org for this is global, place them with staff in opposing time-zones to optimize unnecessary confirmations (though you use a valuable measurement on how consistent your staff is)
> Services that result in the power down of resources will no longer automatically take action on any account, regardless of lack of payment history, for accounts that were engaged more than 90 days prior. These cases will be escalated for manual review.
I can't count how many services I've deployed that started under 90-days ago where the customer failed to add their account information to the service. I can't count them because I don't know. Usually our customer creates their account with instructions from us, and creates an account for us to use which doesn't have permissions over payment details. I wouldn't be surprised if I've had an app go past production that the customer simply forgot to do that important step on Day-1, or if the customer procrastinated until production, etc. We ask, but I've been lied to about stupider things (thankfully rare, but surprising from people who otherwise look like "grown-ups").
Minimally, it sounds like the whole process here is missing a "Hey, WTF is that thing you're running? Call us or we'll need to turn it all off" alert at least a little while before it ... turns it all off. At login, put a clear notice "We want you to love our services, so we let you try them without asking you for payment information. Unfortunately, we have to have monitoring in place to prevent hostile actors from loving us, too. Because of this, accounts newer than 90-days might have services shut off in error. If you want a notification an hour before action will be taken, provide your mobile phone number and we'll send you a text. Or you can enter credit card information/confirm your identity (not sure what options are available here) and we'll keep the bots from bothering you"
Of course, all of this costs money. And based on the incident response times, an explanation other than "failure to prioritize correctly" might very well be "failure to staff properly/have the tooling in place to handle the volume". Considering the competition in this market, I wouldn't be terribly surprised if "we can't afford it" plays into some of that.
[0] A little less awful than AWS in a lot of ways for the task I had to do.
There are, however, a few things that could be improved about this process:
> Peer review of account terminations. For any account appealing a lock, two agents will be required to review the submission prior to issuing a final deny.
The devil is in the details. Do this in a manner that the person confirming that the account is committing fraud is unaware that they are confirming another's denial; otherwise "dude, can you approve that termination I just did? I want it out of my queue/that guy was a dick." is a risk.
Off the top of my head, I'd probably generate two support tickets (linked, but without that link presented to the account termination CSR team member), assigned directly to this person, hidden from others ("hiding", along with training/process improvements is likely enough). If one person disagrees with the termination, close out the other person's ticket. If the CSR sub-org for this is global, place them with staff in opposing time-zones to optimize unnecessary confirmations (though you use a valuable measurement on how consistent your staff is)
> Services that result in the power down of resources will no longer automatically take action on any account, regardless of lack of payment history, for accounts that were engaged more than 90 days prior. These cases will be escalated for manual review.
I can't count how many services I've deployed that started under 90-days ago where the customer failed to add their account information to the service. I can't count them because I don't know. Usually our customer creates their account with instructions from us, and creates an account for us to use which doesn't have permissions over payment details. I wouldn't be surprised if I've had an app go past production that the customer simply forgot to do that important step on Day-1, or if the customer procrastinated until production, etc. We ask, but I've been lied to about stupider things (thankfully rare, but surprising from people who otherwise look like "grown-ups").
Minimally, it sounds like the whole process here is missing a "Hey, WTF is that thing you're running? Call us or we'll need to turn it all off" alert at least a little while before it ... turns it all off. At login, put a clear notice "We want you to love our services, so we let you try them without asking you for payment information. Unfortunately, we have to have monitoring in place to prevent hostile actors from loving us, too. Because of this, accounts newer than 90-days might have services shut off in error. If you want a notification an hour before action will be taken, provide your mobile phone number and we'll send you a text. Or you can enter credit card information/confirm your identity (not sure what options are available here) and we'll keep the bots from bothering you"
Of course, all of this costs money. And based on the incident response times, an explanation other than "failure to prioritize correctly" might very well be "failure to staff properly/have the tooling in place to handle the volume". Considering the competition in this market, I wouldn't be terribly surprised if "we can't afford it" plays into some of that.
[0] A little less awful than AWS in a lot of ways for the task I had to do.