This message was deleted.
# _general
s
This message was deleted.
a
We are thinking if the disk is going to be a problem, either we need dedicated something, which is stable, for fslogix, or we need to think about designing Citrix UPM profiles again.
(The best solution would be for our disk team to fix the underlying disk issues, I know 🙂 )
j
@Jarian Gibson another war story to add to the list. @ATWVD I have not hit this one specifically, but that sounds horrible. I have seen the occasional corruption, but not at this scale. We have forever had to reiterate the importance of stable storage for Container Technology, even with Cloud Cache in the mix, it's there to deal with short term loss - if storage is a concern and you don't have a 100% need for Container tech, then CPM is a way better solution, far more resilient and far less dependent on high end storage. It's a balance, only you know what your unique needs are for container tech vs Files. Note that CPM can also front end FSLogix with file cache to help, but I haven't put this into prod as yet
j
Storage issues can corrupt file based profiles as well.
Less resilient on high end storage but still can get corrupted with storage issues.
j
True, but if the profile is sitting on the VDA and the storage goes kaboom, far less impacting
and UPM has self based healing for that scenario too
j
While logged on yes. During login or logoff… 😬
😬 2
Also depends on features being used.
Still need to address storage either way. Especially resilience and stability.
a
I would prefer if FsLogix could affect ONLY the users with corrupt profiles, instead of hanging the frxsvc on the whole VM and also causing timeouts on other services. Would be easier to cleanup individual users, or have some sort of way to identify the corrupt profiles. Analytics service which could collect all logs and identify this
l
I have seen this exact issue in a very specific scenario. Cloud Cache creates a lock file to specify which machine has access to the disk, this in theory is fine as when a new user logs on the file is updated with the name of the new machine unless the VHD is locked. What I have seen is a lock file is in existence when a VHD file has been deleted. This basically ends up in a locked scenario. Why? The session itself is still running (because the storage failed) and the VHD is detached. The lock file is checked regularly by the cloud cache service, though and then re-created. You essentially end up in a completely locked state which sessions that actually cannot log off. My recommendation. Reboot absolutely everything and remove all the .lock files on mass.
🙌 2
a
Thank you, we will keep this in mind if we ever get a mass corruption incident like this again.
l
Check your FSLogix logs and see what they say during a login. If you're not sure just post some snippets here.
g
I've had at a lot off issues with .lock files as well, in VDI scenarios using pooled/random VDIs. If the VDI for some reason is unexpectedly shut down, the lock file is left behind. When the user logs back in again he/she is most likely routed to a different VDI. The login fails, with FSlogix stating that the profile is in use/locked. Deleting the lock-file almost alwsays solved the issue, but it is a PITA to have to delete file manually each time. The vhd itself is left untouched.