This project proposes a Modified U-Net (M-U-Net) architecture for multi-source audio stem separation, aiming to extract vocals, drums, bass, and other instruments from a single mixed track. Unlike traditional methods that train separate models for each source and often favor louder instruments, M-U-Net uses a unified encoder-decoder structure with skip connections and enhanced loss functions (Dynamic Weighted Average and Energy-Based Weighting) to balance separation quality across sources. Trained on the MUSDB18 dataset, the model demonstrates competitive performance with fewer parameters and faster inference compared to state-of-the-art approaches. The system design includes preprocessing, model training, inference, and user interface integration, with applications in remixing, karaoke, audio restoration, and music information retrieval.