I spent a year as a High Performance Computing Sys Admin. After careful consideration, we just decided reliability wasn't worth the cost. We bought more hardware instead of a UPS system for our data center. Once or twice a year, we experience a power blip and lose all our compute nodes (infrastructure is on UPS). We would send out an apology email to our users and tell them to resubmit their jobs. Worst case scenario was someone had to restart a 14 day job. Had we bought the UPS system, that 14 day job would take almost twice as long to complete, spending a week or so longer in the queue. We started embracing the idea of acceptable failures and saw it had a great impact on our service.
Do your users run 14day jobs with no checkpointing whatsoever? I'd be afraid of a bug in my code crashing the computation 90% of the way through. The MapReduce setup seemed much more resilient to things like this for example.