Those sizes are rules of thumb ... I have seen segments up to several GB in size, and upwards of 60m rows.
And in my experience the two things you want to monitor to make sure the segments aren't too large are:
• Peon scan times for streaming ingestion, for when the data is still held in the streaming ingestion task. Peons are generally slower than Historicals for conducting segment scans (their data is more fragmented, and some of it is unindexed), so it can help to reduce Peon segment scan time by building smaller segments more frequently
• build/publish times during ingestion when the segments are first created ... "normal" segments take 3-5 minutes to build and publish ... but larger segments can take up to 10-20 minutes or more. You just need to make sure this doesn't interfere with your ingestion task operations.
• Historical segment scan times will tell you if the segments are too large. For high QPS good scan times are in the single digit ms ... for low QPS you can probably tolerate up to 300ms scan time. However the scan time is also highly dependent on the query being executed.
Thanks. John