A robust, flexible voice cloning system with advanced speaker diarization and configurable resource management.
- Advanced speaker diarization
- Flexible resource allocation (CPU/GPU)
- Comprehensive test coverage
- Efficient caching system
- Modular pipeline design
- Python 3.13.2
- PyTorch 2.1.0
- torchaudio 2.1.0
- Compatible with both CPU and GPU environments
- Clone the repository:
git clone [your-repo-url]
cd voice-clone-engine- Install dependencies:
pip install -r requirements.txt🚧 This project is currently under development. For development status and planned features, see:
REFACTOR_PLAN.md- Detailed development roadmapREQUIREMENTS.md- Project specifications and requirements
The system supports flexible resource allocation:
from src.resource_manager import ResourceManager
with ResourceManager() as rm:
# Get available resource information
resource_info = rm.get_resource_info()
# Configure resource usage
rm.configure("--resource all") # Use all available resources
rm.configure("--resource cpu") # CPU only
rm.configure("--resource gpu") # GPU only
rm.configure("--resource count:3") # Use any 3 available resources
rm.configure("--resource-exclude cpu") # Use all except CPU
rm.configure("--resource-memory-limit 80") # Limit memory usage to 80%
rm.configure("--resource-strategy balanced") # Distribute load evenlyThe project includes several demos:
cd demos
python embedding_demo.py # Voice embedding visualization
python diarization_demo.py # Speaker diarization
python caching_demo.py # Performance impact of cachingvoice-clone-engine/
├── configs/ # Configuration files
├── data/ # Dataset storage (gitignored)
├── demos/ # Demo scripts
├── src/ # Source code
│ ├── speaker_diarizer.py
│ └── resource_manager.py
├── tests/ # Unit tests
└── plots/ # Visualization outputs
Run the test suite:
python -m pytest tests/For development guidelines and testing practices, see REFACTOR_PLAN.md.
[Your chosen license]
See REFACTOR_PLAN.md for development guidelines and workflow.