At Amazon's scale, a very very small increase in the effectiveness of recommendations can represent millions in revenue. They probably do have layers of exceptions, and custom-built software used by teams of "editors" to manually manage these exceptions. Not to mention the experimentation system where they are simultaneously testing hundreds of things from CSS one-liners to new ranking algorithms.
It's probably a pretty complex beast with the weight and inertia of a core system that's been continuously in production through 20 years of growth. They probably continue to invest heavily in it, the ROI is there, it's just a really hard problem for them at this point. Their engineers probably curse its limitations and would love to rewrite/replace old and outdated parts, and indeed there will be teams working on that, but those projects will either die out or spend so much time reaching feature parity with the bloated existing system that it ends up looking much the same.
It's probably a pretty complex beast with the weight and inertia of a core system that's been continuously in production through 20 years of growth. They probably continue to invest heavily in it, the ROI is there, it's just a really hard problem for them at this point. Their engineers probably curse its limitations and would love to rewrite/replace old and outdated parts, and indeed there will be teams working on that, but those projects will either die out or spend so much time reaching feature parity with the bloated existing system that it ends up looking much the same.