Which statistical method is most appropriate for comparing performance rates across hospitals when sample sizes vary significantly?