Stereonetworks trained on synthetic data often degrade in real-world scenes. Within an IGEV-style pipeline, domain-sensitive features may disturb correspondence cues, average matching-dimension pooling may attenuate reliable responses, and local motion encoding does not explicitly exploit rectified stereo geometry. We propose Evidence Preservation and Propagation Stereo (EPP-Stereo), which coordinates feature refinement, volume construction, and recurrent aggregation. Frequency-decoupled refinement regulates low- and high-frequency cues; a stereo-oriented Peak-Preserving Volume Pyramid applies channel-wise temperature-scaled softmax weighting to adjacent disparity responses in both the geometry-encoding and all-pairs correlation pyramids; and Epipolar Large-Kernel Aggregation injects local, epipolar, and orthogonal contexts before the original ConvGRU. Component and directional-kernel ablations further support the complementary roles of the proposed modules. Although the resulting design requires additional computation, EPP-Stereo, trained only on SceneFlow, reduces zero-shot Bad 2.0 on Middlebury from 7.1% to 6.0% at half resolution and from 6.2% to 5.1% at quarter resolution, and lowers Bad 1.0 on ETH3D from 3.6% to 3.4% relative to IGEV-Stereo.
