Hi,There are a few ways to get around this depending on the type of audio you are working on.
If you are using a piece of pre-recorded audio with two characters talking then you can separate the two voices with an audio editor such as the free Audacity. Then the timing will be spot on when you add each voice to the characters.
If you are using text to speech voices or recording your own voice(s) then you can break the script up into smaller parts and work with these. For example in a two person conversation, you could record character ones first bit of speech, then save this. Then add the second persons response and so on. This way you will have many sections of script in Stage which you can position perfectly to get a good flowing conversation.
You can also simply jot down the times of the gaps in speech when previewing character one and use this as a guide when working with character two. This will work as long as there isn't too much fast interaction going on between the characters.