Abstract
We present a two-stage deep convolutional neural network for human pose estimation. In the fi rst stage, it directly extracts features from the input image and combines all the features to generate a compact yet e ffective result for predicting the keypoint locations instead of producing one heatmap for each keypoint. Then, we use the input image and the synthetic heatmaps derived from the previous stage as the input of the second stage to get a refi ned result of pose estimation. We evaluate our method on two datasets: FLIC and LSP. Our method achieves the state-of-the-art performance on FLIC dataset.