This message was deleted.
# citrix-netscaler
s
This message was deleted.
j
Hi John, I posted this same error a couple of weeks back. New deployment for moving from RADIUS to Azure AD/SAML. only one user has seen it a handful of times and as recent as yesterday. Of course the user is the infosec manager and this has caused a bit of uncertainty. I opened with Citrix and they said its very odd as the would expect this error to occur all the time vrs intermittently but never really got/found anything solid, cheers
c
Thanks m8 keep us updated. The session profile has the IP of the SF LB VIP so not DNS and i dont see the SF LB VIP flapping at all when i checked the newnslog
May 19 111152 <local0.info> 10.129.75.211 05/19/20231011:52 GMT IFSRNETS01 0-PPE-1 : default SSLVPN Message 5223371 0 : "AAA Client Handler: Found extended error code 917511, ReqType 16388 request /cgi/logout" May 19 114719 <local0.info> 10.129.75.211 05/19/20231047:19 GMT IFSRNETS01 0-PPE-0 : default SSLVPN Message 4971822 0 : "AAA Client Handler: Found extended error code 917511, ReqType 16388 request /cgi/logout" May 19 115311 <local0.info> 10.129.75.211 05/19/20231053:11 GMT IFSRNETS01 0-PPE-1 : default SSLVPN Message 5260575 0 : "AAA Client Handler: Found extended error code 917511, ReqType 16388 request /cgi/logout"
few error's in the ns.log
https://twitter.com/rsrevord/status/1580281657956847616?s=20 @RSRevord did you ever find the cause was it the firmware ? sounds like it was a memory leak ?
👍 1
r
trying to remember give me a few
this was AAD not okta, but here is one for okta on this specific error https://support.citrix.com/article/CTX224410/http11-internal-server-error-43524-after-oauth-authentication
c
Thanks Yeah seen that ctx article earlier it's Oauth tho not SAML so is different
j
@RSRevord Thanks bud, @c4rm0 any other ideas based on Ryans feedback? My customer has now put a hold on this a delayed the rollout til I have a solution, so frustrating! 😞
r
in your SAML what do you have for logout binding? Redirect or Post?
c
It's set to post
r
use redirect
👍 1
c
Will give it a try
r
also pretty sure you can recreate this pretty easily, its people never closing their browsers post logoff, if they open in private the problem goes away, or reboot
i might have a couple other work arounds as well
but as i recall redirect was the money
j
My logout binding was already on Redirect, SAML Binding on POST:
c
@John Gallacher Do you use the metadata xml (import metadata) to configure your SAML auth policy or you got it manually configured??
j
Manually bud, the import checkbox is unchecked, ta
c
mine is manual on this implementation as well (I norm use the import metadata option)
j
Ive had about 10 folks testing for a good few weeks.. maybe close to a month. One of the project stakeholders has seen the issue around 3 times in total. Thats it. Im worried Im chasing a ghost here and its a client / browser issue? cheers
c
Someone messaged me other night (IT guy) who had same error I didn't think much of it but then I got the same error myself but have been unable to reproduce it and no one else has had it. As Ryan said I think it's a browser or client issue where the browser isn't closed after logging off
r
login let the browser sit and until it times out etc, then wait another good 30 i bet you get the error
its stale cookies
least my theory, i had another work around to lower cookie i think it was give me a few another team mate suggested and tried that I believe too many issues rolling around in my head to track
j
Yeah they have said when the refresh the page or hut the URL again (from the same session) its good and they get to SF...
I will have a look at the cookie timeout settings..
k
in some cases I have ended up creating a responder that hits if a certain cookie exists and the requested url matches... then with a redirect action just wipe out the cookie and get it sorted out
it's pretty ugly and shouldn't be like that though
j
Thanks @Kari Ruissalo (WyW). My LBs etc are all set to source IP vs Cookie insert/timeout. Is there somewhere else I should be checking for this? cheers
I just got it for the first time logging in this AM, refreshed and got a follow up error.. removing saml part of the url and hitting again got me straight onto SF. I will check the logs and will feedback, cheers
@c4rm0 do you have anything set in here?
c
@John Gallacher nope got nothing set on authentication host or authentication domain I never norm set those. I only started getting this issue since upgrading to 13.1-45.63 when I was 12.1 never had the issue
j
Yeah same, never seen it before. Im now just clutching at straws, might get Citrix back on the blower..
@c4rm0 reopened this with Citrix and shared a copy of this thread. I'll keep you posted. Out of interest, do you have GSLB in play here? Reason I ask is that I do, Ive only seen the issue first hand since removing a hostfile entry that was pointing me a the main site, cheers
c
Yeah using GSLB
j
Citrix have just told me they believe this issue is fixed in a new FW... for 13.0 they are saying 91.12nc. They also asked me to increase the SAML skew time from the default 5 minutes to 15 minutes. We have eG innovations in place, we are also going to setup a monitor/alert for the error in question. They said if the upgrade was going to take some time to get done the following command could be run as a workaround but I havent delved into what its actually doing yet, command is: nsapimgr_wr.sh -ys call=ns_aaa_saml_disable_context, cheers
Does anyone know what Netscaler log these http/1/1 internal server error 43524 would land in? Trying to set up eG to capture. Just got one report of it happening this AM but its not been caught. I have this set:
c
Guessing it would be either in the nslog or one of the http logs like httpaccess-vpn.log or httpaccess.log Did you find out any info on that nsapi command Citrix recommended ? I can't find any info on it at all
j
Cheers, John. I will ask them for more info and will feedback today bud.
f
Hi, I changed the SAML skew time from the default 5 minutes to 15 minutes, No issue reported on my VPX since.
j
Nice... it hasnt worked for me unfortunately. FW upgrades taking place tomorrow evening, cheers
k
when you're facing the issue, have you tried to see what is caught in ns.log so that why the SAML fails?
tail -f /var/ns.log | grep {uuid}
where the uuid is the username like "first.last@company.com" ... also change the AAA authentication log parameters for the test to debug if the normal level doesn't reveal enough?
j
Nice one @Kari Ruissalo (WyW). I'll give that a try...
@c4rm0 Got this back from Citrix on the cmd: "nsapimgr_wr.sh -ys call=ns_aaa_saml_disable_context is used when the SSL VPN Gateway rejects client-side SAML connections"
k
I don't love those "under the hood" knobs that do "something"
🙃 1
j
Man, more reports of this after it being quiet for a week since doing the FW upgrade 😞 IVe gone ahead and run nsapimgr_wr.sh -ys call=ns_aaa_saml_disable_context while on with Citrix there, now clutching at straws... they also questioned token lifetime/expiry etc so I will check with the client as ive not got access to Azure AD/DUO, cheers
c
yeah i am getting more and more reports of the issue. I logged it with support so will update the thread as well. It does seem to only happen if they leave their browser open for a long period (overnight ect) i will also look at increasing the SAML Skew time as its set to 5 mins
I managed to capture the error in the ns.log and sent it to citrix and they recommended the same nsapimgr command Jun 21 132131 <local0.info> 10.129.75.211 06/21/20231221:31 GMT IFSRNETS01 0-PPE-0 : default AAA Message 17230054 0 : "nFactor: deserialize aaa_info, timestamp verification failed" Jun 21 132131 <local0.info> 10.129.75.211 06/21/20231221:31 GMT IFSRNETS01 0-PPE-0 : default AAA Message 17230055 0 : "SAML deserialize error: failed to extract saml action or failed context" The solution to the problem is the following : shell command: nsapimgr_wr.sh -ys call=ns_aaa_saml_disable_context
Have you seen it since applying it? @John Gallacher
j
Thanks @c4rm0. I struggled to capture the error but with that said, I was still getting it post fw update (I got told the cmd wasnt needed). I had another call with them last week and they told me to run it anyway (and reboot). Its now in place and the reboots were done on Tue evening. Awaiting feedback 🙂
It was happening at least 3 or 4 times per week.. lets see, fingers crossed 🙂
👍 1
I forgot I replied last week... the reboots only happened this week which I got told were needed.. Ive had no reports since last week... so its promising.
c
i asked citrix if it the nsapimgr command persists across reboots and they said it did which was strange as you normally have to add it to the rc.netscaler file which i have done. After me questioning them they have come back said it does need to be added to the rc.netscaler file for it to persist
j
FFS lol... that explains it. I just for more reports of it today...
Ok, cmd added to the rc.netscaler file. I'll need to reboot again for that config to load right? ta
@c4rm0 - Ive just replied to Citrix to question this too. They should have known this 😞
c
If the nsapimgr command has already been run via cli you shouldn't have to reboot for it to take effect. Adding it to the rc .netscaler file allows the setting to be retained after the reboot as it would revert to default setting.
j
Nice one.. Its been run again and added to the rc file.. lets see how we go 🙂 😅
c
Well i ran the command nsapimgr_wr.sh -ys call=ns_aaa_saml_disable_context and added it to the rc.netscaler file but didnt reboot and had reports of same issue this morning so guess it does need a reboot (Some nsapimgr command i ran in past dont need a reboot) will see how we get on after the reboot
j
Yeah.. I will get my deployment bounced on the weekend
c
I not had any reports of issues since applying it
j
@c4rm0 - thanks mate, I was out on holiday the last 2 weeks, I will check in and see what the status is. Sounds promising based on your feedback, cheers
@c4rm0 - Hey mate, confirming Ive had no more occurrences since adding the cmd/config to the RC file. The project was put on hold, all looking good now. I still have random RPC errors on the FAS servers for cert failures. There are 4 FAS servers and these errors dont seem to be user impacting (i can see my name in there too). Once I clear these and work out whats causing I should finally be able to schedule a go live... what a journey!