Showing posts with label data center profiling. Show all posts
Showing posts with label data center profiling. Show all posts

Tuesday, September 22, 2009

Making cloud computing work: customers at 451 Group summit say costs, trust, and people issues are key

A few weeks back, the 451 Group held a short-but-sweet Infrastructure Computing for the Enterprise (ICE) Summit to discuss "cloud computing in context." Their analysts, some vendors, and some actual customers each gave their own perspective on how the move to cloud computing is going -- and even what's keeping it from going.

The customers especially (as you might expect) came up with some interesting commentary. I'm always eager to dig into customer feedback on cloud computing successes and roadblocks, and thought some of the tidbits we heard at the event were worth recounting here.

(A side note: if you're interested in more cloud-related customer comments, you can look at some previous Data Center Dialog posts, including this one recounting questions overheard about internal clouds from a few months back.)

Clouds under the radar

As a way to set the stage, 451 Group analyst and ICE practice lead Rachel Chalmers compared cloud computing’s adoption to that of Linux in the late '90s, and to the beginnings of server virtualization in dev/test environments. "There was a lot of VMware [being used in IT] before CIOs even knew it was there. It only belatedly comes to the attention of the architects." Chalmers was clear that adoption is underway ("Customers are already using cloud," she said), and emphasized how under-the-radar adoption like this can really help. "They come pre-evangelized," said Chalmers. That means there are a lot fewer people to convince when it comes time to make the case to roll this out widely.

Customers: Some hesitate to call it cloud

Yuvi Kochar, CTO from the Washington Post Company, saw the work that his teams were doing as shared services, and said, "I hesitate to call this work a cloud." He acknowledged that they were enabling elasticity and cost accounting, but that some of what they were working on was still "hand-wired." What was the thing that pushed Kochar over the edge to move toward dynamic, shared IT services (even if he can't bring himself to actually it "cloud")? Cost. "We want to move everything to a variable cost model," Kochar said. Another customer, Ljubomir Buturovic, VP & chief scientist from Pathworks Diagnostic, said they told their server vendor they would be buying no more hardware, and they haven't. Again: cost was the driver -- in this case, capital costs.

Cloud: It's (still) not for the faint of heart

I thought some of the most useful insights of the customer panel were from Jim Houghton, now co-founder and CTO of Adaptivity. Houghton reflected back on his experiences a few years ago getting what could now be termed a private cloud -- then called utility computing -- off the ground at Wachovia and Bank of America. He characterized the work as something which eventually turned into a $100-million project to stabilize a "chaotic environment" in IT and "took out $500 million in op ex."

Houghton noted that he and his employers at the time "learned a lot of hard lessons along the way." For example, the mechanisms for truly automated provisioning and handling dynamic shifts in demand are "really important," but really hard, especially since the tools at that time he started (early to mid-2000s) were still quite immature. Metering what you're using is critical, he said, but "it's hard to meter when you're separate from the physical environment." Overall, said Houghton, cloud-style dynamic infrastructure is "not for the faint of heart."

Funny. I heard Donna Scott of Gartner make the same observation about creating a real-time infrastructure back in December. We've come a long way, but there's more to do, for sure.

Biggest pain: impact on the people and the organization

Houghton identified a couple best practices for making the shift to a dynamic, shared infrastructure to support your applications: for example, workloads that are entirely self-contained and in which you have access to the source code make excellent candidates for these types of deployments. However, he said, you need to truly understand the workload characteristics. That last bit is something I've heard many large IT organizations lament as a huge problem -- profiling what they currently have and how it runs.

However, he said, the more painful thing is the change in the operational model and its impact on a company's organization and the people that work in it. I've heard this many times over the past several years, from customers, analysts, and even vendors. You can't underestimate the impact that the cloud operating model will have on personnel.

"It gets to be a very touchy subject for clients," said Houghton. "It's hard to do the business cost model without talking about 'these 20 people will be out of a job.'" Or at least who can be moved to focus different things than they're focused on today.

Need to move beyond just virtualization

Houghton also made a comment that infrastructure that's "mildly virtualized may be efficient, but is not all that dynamic." I have to agree: there is a lot more needed to make a private cloud-style infrastructure fly than loading up on a bunch of virtual servers. In fact, that point was made pretty strongly when Dan Kusnetzky lined up his fellow 451 Group analysts to highlight some of the cloud computing inhibitors. William Fellows conveniently rattled off some of his (and the industry's) favorites: security, SLA support, corporate governance, interoperability, vendor viability, job security, and misaligned business models. To name a few.

These are issues we've heard before, for sure, and are at the top of the list of things to be addressed in order for cloud computing to be viable day-to-day in an IT shop. I think it's good news that we're hearing more and more about the manageability side of the issue.

Can I drive your Mercedes while you're not using it?

One of the strongest objections to an internal cloud, or really a shared infrastructure of any sort, still boils down to what's called "server hugging." Houghton gave an amusing explanation of the mentality by putting it this way: "Just because I have 4 Mercedes and I can only drive 1 at a time, doesn’t mean I'm going to let you drive the other 3." In his case, the team of coworkers and vendors pitching this new approach had to put in "years of work" to "build up the trust. What you have to say in response is, 'You'll get everything you wanted, plus a lot more.'" And, of course, you have to back it up by delivering on your promises with great cost savings, excellent service levels, and an improved ability to respond to new requirements from the business.

Are we making progress on cloud computing?

Chalmers noted that in many cases a move to cloud computing doesn't feel like progress at all. "All we're doing is moving the headaches somewhere else. But," she said, those management headaches "still need to be solved."

Fellows noted that there really isn't any definition of success so far. Early adopters of grid, utility computing, and virtualization have been the ones in his experience to be most aggressive in working toward cloud-style environments. "It’s a logical end-point for any of those [earlier] activities," said Fellows.

In fact, said Chalmers, "very often when we see an early adopter of cloud that's successful, it's because they understood HPC [high-performance computing] and putting everything under control in the data center."

So, are we making progress toward incorporating cloud computing in today's IT environment?

Houghton from Adaptivity made the point that "the economic malaise has put a lot of power back into the CTO's hands." It's a chance to use this power to instigate more sweeping changes to how IT operates than any time in recent memory. But people are being judicious with that power.

Appropriately, Chalmers probably did the best job of putting cloud computing in that context: "In this world of financial crisis, the acid test of any technology is: 'Does anyone care enough to sign a purchase order?'" Clearly, some do. (And some fraction of those are on stage at events like this talking about what they've done and learned.) And, as the industry matures the management capabilities and works out the kinks that many of these customers noted, others will start to feel that it's time to sign on the dotted line, too. And hopefully join them on stage.

Tuesday, March 31, 2009

No April Fools' Day joke: Data center managers don't know what their servers are doing

Given the arrival of my favorite Silicon Valley holiday, I'd like to brush aside some of the content in this post as a big April Fools' Day ruse. Unfortunately, it's not.

Here are the facts: according to a couple folks who should know, data centers are buying new equipment before making good use of what they already have. OK, maybe that's not new news, but we've just added another couple scary bits of data ourselves. According to a new survey we did here at Cassatt, not only are data center managers making poor use of their equipment, they don't even know what some of that equipment is doing.

Sounds like a major disconnect. Here are some specifics from a couple big analyst firms and early feedback from our 2009 Cassatt Data Center Survey:

Gartner: Organizations struggle to quantify their data center capacity problems

Gartner analyst Rakesh Kumar published a paper (ID #G00165501) at the beginning of the month in which he slapped some pretty direct zingers (for Gartner, anyway) in his "key findings":

· 50% of data centers will face power, cooling, and floor space constraints within 3 years. (Our forthcoming survey, by the way, saw similar problems. 46% said their data center is within 25% of its maximum power capacity.)
· Data centers use "inefficient, ad hoc approaches" rather than "continuous-improvement, process-driven" methods to get a handle on these problems (how's that for a nice way to scold the IT guys?).
· Most organizations can't quantify their capacity problems.

Now, we know that the processes, like the technology, in use in today's data centers (especially for big organizations) have been cobbled together over time. Those ad hoc approaches Rakesh mentioned don't surprise me. But not being able to even quantify the problem seems like an issue that has to get solved immediately.

Forrester: the bad economy means it's time to improve IT operations...or else

Forrester's Glenn O'Donnell published some similar points in a recent NetworkWorld article. Given the rocky economy, companies are making IT investments in anything that improves operational discipline, provided there are speedy results, according to Glenn. He places the operational IT budget at around 75% of the total that companies spend on IT (also known as the "keeping the lights on" money). However, "30% to 50% of the energy we expend is wasted," he said, blaming "inefficient processes and poor decisions." And by energy, he means time and effort. And time = money. So he's talking about...money. The business world equivalent of Darwin's natural selection will not be kind to companies wasting money in the current economic climate, says Glenn. He points to "mean time to resolution" (MTTR) as a gauge of how a company is doing in its data center operational efficiency. One solution: "We must demonstrate...MTTR improvements with pilots of process and automation."

Cassatt survey: people don't know what their servers are doing

Back for a moment to Rakesh's Gartner report. "Most organizations," Rakesh writes, "struggle with quantifying the scale and technical nature of their data center capacity problems because of organizational problems, and because of a lack of available information." There's the rub. You need real, actual data before you can do anything about it. And -- as our customers have been telling us -- that's not easy to come by.

Our new Cassatt 2009 Data Center Survey (due out in the next few weeks), shows how acute the problem is. The survey will show that over 75% of data center managers only have a general idea of the current dynamic usage profile of their servers. A couple other somewhat disturbing stats we found:

· 7% said they don't have a very good handle on what their servers are doing
· 20% know what their servers were originally provisioned to do, but aren't certain that those machines are actually still involved in those tasks
· Only a bit more than 16% of those in IT ops have a detailed, minute-by-minute profile of what activity is being performed, the users involved, granular usage stats, interdependencies, and the like
· More than 20% of respondents thought that between 10-30% of their servers were "orphans" (servers that are on, but doing absolutely nothing). The actual number we have routinely found to be true orphans in our investigations with big customers, incidentally, is right around 11%. (For more on the "orphan" server topic, see my previous post.)

Getting the information that IT ops needs

OK, so that's all pretty dire. I'm interested, though, in how we help end user IT ops teams make progress in the face of this. I know that when working with customers on data center efficiency projects, this lack of data doesn't cut it: it's much better to come to them with a solution. Or at least some suggestions.

I'm sure others have come up with different ways to solve this, too, but we at Cassatt realized we had to create a way to get a profile of what a customer's environment is actually doing over time. So, as a step in the process of using our software to create an internal compute cloud -- complete with automated service-level management policies for their applications, VMs, and servers -- we put together an Active Profiling Service. We use some monitoring software and tap into the smarts from some of our experts who know the ins and outs of data centers to put this service together. The result: a look at what your data center is doing today and recommendations about what to do with that info.

(If you're interested, we can show you some sample Cassatt Active Profiling Service reports: ping us at info@cassatt.com. Also, Steve Oberlin and Craig Vosburgh will be walking through some aspects of this in a webcast this week.)

Once you have the data: some data center optimization suggestions

Once you have some of the crucial profile data, what are some useful optimization suggestions?

Glenn from Forrester points to Harley Davidson as someone who is headed in the right direction through a combination of "process, automation, hiring good people, and a determination to discard the destructive practices of the past."

Randy Ortiz of Data Center Journal suggests a little something called the "Data Center and IT Ops Diet." When you are on a diet, Randy notes, "you carefully examine what you take in and how much you burn off. The Data Center and IT Ops diet [which he describes as driven by the economic downturn] provides you with the necessities only: availability and efficiency. There is no room for large projects with long-reaching ROIs."

And what about Rakesh of Gartner? He has similar advice: before jumping into some new data center build-out, a "continuous process of data center improvement [should] be established" (in fact, his whole paper is called, appropriately enough, "Continuously Optimize Your Data Center Capacity Before Building or Buying More," which is what got me started on this whole rant in the first place). "Too often," Rakesh writes, "the data center improvements are considered a project" and not a long-term, on-going process. He suggests continually optimizing IT infrastructure because "sprawl in the infrastructure creates sprawl in the data center." The tactics he lists for consideration are pretty basic: consolidation, virtualization, and tossing out older hardware.

These are great starts. We're seeing customers do them all. And many of those customers we are working with are implementing these ideas as an integral part of a data center optimization project that also includes working with our software and an active profiling engagement (we do, of course, as a result, suffer from a sampling bias). However, what we've been helping these customers do is relevant here: they can identify and decommission specific orphan servers -- the actual ones that are sitting around doing nothing. They can find the specific candidates for virtualization, where server utilization is very low, and workloads are such that they can be stacked together with other workloads on fewer boxes. They can locate and set up intelligent server power management -- identifying hardware for which workloads are very cyclical, enabling servers to be completely shut down to save power during off hours. And, from all this can also come recommendations for ways to optimize your overall IT operations, including setting up policy-based IT infrastructure automation and an internal compute cloud. You could even use some of the decommissioned servers as spare capacity.

As cool as any of this may sound, though, there's no need to get carried away. Start simply. Find out what your servers are doing. Or start experimenting with optimization in a limited corner of your environment where the stakes aren't very high.

But start.

The industry that spawns some of the most creative April Fools' Day jokes shouldn't be one itself.

If you're interested in getting a preview copy of our second annual Data Center Survey results under non-disclosure, let me know at jay.fry@cassatt.com.

Tuesday, December 16, 2008

Killing comatose servers: OK, but how?

One of the best sources of data center energy efficiency guidance available today is the Uptime Institute. Not only do they run some really focused, useful events on the topic, but their fearless leader, Ken Brill, is very visible and very direct with his recommendations. His recent article in Forbes took on one of those dirty little secrets in IT: there are a lot of servers in your data center doing absolutely nothing.

Given how serious of a problem that the current growth rates of data center energy usage will be, Ken and Uptime have given some serious thinking to how to curb the problem. Topping that list was the directive that served as the title of his article: "Kill comatose computers."

Call them orphan servers, comatose servers, idle servers, or whatever, Brill calls them "corporate enemy No. 1. Unless you have a rigorous program of removing obsolete servers at the end of their lifecycle," he writes, "it is very likely that between 15% and 30% of the equipment running in your data center is comatose. It consumes electricity without doing any computing."

Yikes. Numbers tossed around by Paul McGuckin at the recent Gartner Data Center Conference are similarly high. The solution? Brill has a simple answer: "This dead equipment needs to be unplugged and removed."

No arguments so far from me. In fact, doing work like this is one of the steps we recommend toward revamping and improving how you run your data center. The more intriguing question is one Ken also asks: why hasn’t this already happened?

The answer, unfortunately, is that most shops have worked long and hard on the steps for standing up servers, installing new components, and the like, probably because adding things to the data center is always done with some sort of time or business pressure. People are watching and they want their stuff up now. Rarely is someone breathing down your neck to unplug something. In the frenetic everyday life of an IT ops person, the decommissioning bit is the part that can wait while you handle the urgent fire of the day.

The problem is that after today's crisis comes tomorrow's. And though the orphan servers are using up power to keep them running and air conditioning to keep them cool, removing them isn't a priority. But as more and more data centers start to hit the wall for power capacity, that's going to have to change.

So, what do you do? As Ken Brill points out in Forbes, "after weeks or months pass and employees turn over, the details of what can be removed will be forgotten, and it becomes a major research project to identify what is not needed."

Unfortunately, identifying orphan servers is not something that will take care of itself. Here are some of things we've seen customers focus on to help solve this problem:

- Enable some sort of detailed monitoring on your servers
- Determine what time period is appropriate to watch for changes, based on your business and what the servers are likely to be doing
- Watch usage, users, processes, and other statistics that will be helpful in making decisions later
- Sift through the mountains of data you collect with someone who can translate it into useful information
- Engage the end users in the process

We've found that organizations sometimes want to do these steps themselves, sometimes they don't. When customers ask Cassatt to help with these steps, it's often because they are in need of the expertise or tested tools and processes for finding out what their servers are doing. It's something they often don't have internally.

The other thing we generally bring to the process is experience we've had working with some very large customers. Through our recently announced Cassatt Active Profiling Service, we've helped customers identify orphan servers, recommended candidates for virtualization, found candidates for power management, and located other servers that they could start to use as a free pool of resources to support a move toward setting up a sort of "internal cloud" architecture.

I guess that's the good news: with the data you get from a project like this, you can start to really make some significant changes to the way you manage your IT infrastructure. So, not only can you follow Ken Brill's advice and begin to kill off those comatose servers and save yourself a great deal of power, but you can also arm yourself with some unexpectedly useful information.

For example, if you know what your servers are doing (and not doing) at different times of the day, the month, and the quarter, you can use that information to start to set up some automation to manage your infrastructure based on those profiles. You can set up shared services to allocate or pull back servers for the applications which have the highest priority at any given time, based upon your priorities.

But I'm getting ahead of myself. The first step, then, is to find out what the servers in your data centers are doing. Then, if you don't like what they're doing, you can actually do something about it.