Hello guys, found something weird. When I've made ...
# help
a
Hello guys, found something weird. When I've made an API request for this domain
<http://brnw.ch|brnw.ch>
one of the result was coming from this company
<http://jcb.co.uk|jcb.co.uk>
Is this normal ? Because they are not the same company at all
c
Hi @Alexis Girard Thanks for reaching out. This is linked with an open bug we have and it will be solved at the end of next week. cc: @Xoel
a
Hi @Christian, thanks 👋
Am I able to identify companies that are impacted ? I'm a bit worried about importing false data into my CRM
x
Yes, this is probably because JCB used https://www.brandwatch.com/ to add some short domains to their LinkedIn URL for example, and we took the domain of brnw.ch as if it was from JCB. This will be fixed in the upcoming weeks
a
Ok so it should be only one corner case ?
x
That’s not the only URL shortener companies use, if you look for bit.ly or even sites like facebook.com, linktree, etc you’ll find also false positives for the same reason. All will be addressed between this and next week
a
Yes I'm already filtering out all of URL shortener but wasn't aware of brandwatch one
By curiosity, how are you going to handle URL shortener ?
x
Ahhh, cool then 🙂 Yes, it’s the first time I hear about them as well
In the cleanup we’re doing of that process we have a list of regex patterns for url shorteners and also social sites like facebook, linketree, etc. If the URL matches any of the patterns, we’ll discard the domain of the URL of that company
and we merge company info from several sources so even if one URL is ‘bad’, chances are high that we get the right URL from other sources
👍 1
a
Makes sense 👍
Copy code
const genericDomains = [
  "<http://bit.ly|bit.ly>",
  "<http://tinyurl.com|tinyurl.com>",
  "<http://rb.gy|rb.gy>",
  "<http://cutt.ly|cutt.ly>",
  "<http://is.gd|is.gd>",
  "<http://v.gd|v.gd>",
  "<http://soo.gd|soo.gd>",
  "<http://clck.ru|clck.ru>",
  "<http://t.ly|t.ly>",
  "<http://shorte.st|shorte.st>",
  "<http://kutt.it|kutt.it>",
  "<http://u.to|u.to>",
  "<http://tr.im|tr.im>",
  "<http://short.io|short.io>",
  "<http://rebrand.ly|rebrand.ly>",
  "<http://tiny.cc|tiny.cc>",
  "<http://adf.ly|adf.ly>",
  "<http://bc.vc|bc.vc>",
  "<http://linkvertise.com|linkvertise.com>",
  "<http://shorturl.at|shorturl.at>",
  "<http://zi.ma|zi.ma>",
  "<http://mcaf.ee|mcaf.ee>",
  "<http://t2m.io|t2m.io>",
  "<http://clicky.me|clicky.me>",
  "<http://t.co|t.co>",
  "<http://lnkd.in|lnkd.in>",
  "<http://youtu.be|youtu.be>",
  "<http://fb.me|fb.me>",
  "<http://instagr.am|instagr.am>",
  "<http://redd.it|redd.it>",
  "<http://amzn.to|amzn.to>",
  "<http://etsy.me|etsy.me>",
  "<http://pin.it|pin.it>",
  "<http://spoti.fi|spoti.fi>",
  "<http://apple.co|apple.co>",
  "<http://ebay.to|ebay.to>",
  "<http://tiktok.com|tiktok.com>",
  "<http://ow.ly|ow.ly>",
  "<http://hubs.ly|hubs.ly>",
  "<http://hubs.la|hubs.la>",
  "<http://hubs.li|hubs.li>",
  "<http://hubs.lu|hubs.lu>",
  "<http://hub.li|hub.li>",
  "<http://hub.ly|hub.ly>",
  "<http://hub.lu|hub.lu>",
  "<http://linktr.ee|linktr.ee>",
  "<http://lnk.tr.ee|lnk.tr.ee>",
  "bio.link",
  "campsite.bio",
  "<http://taplink.cc|taplink.cc>",
  "many.link",
  "<http://buff.ly|buff.ly>",
  "<http://spr.ly|spr.ly>",
  "<http://zpr.io|zpr.io>",
  "<http://smarturl.it|smarturl.it>",
  "<http://onelink.to|onelink.to>",
  "<http://lnk.to|lnk.to>",
  "<http://dlvr.it|dlvr.it>",
  "<http://shar.es|shar.es>",
  "<http://goo.gl|goo.gl>",
  "<http://j.mp|j.mp>",
  "<http://wp.me|wp.me>",
  "<http://nyti.ms|nyti.ms>",
  "<http://lat.ms|lat.ms>",
  "<http://wapo.st|wapo.st>",
  "<http://bloom.bg|bloom.bg>",
  "<http://cnn.it|cnn.it>",
  "<http://es.pn|es.pn>",
  "<http://for.tn|for.tn>",
  "<http://ti.me|ti.me>",
  "<http://bzfd.it|bzfd.it>",
  "<http://techcr.ch|techcr.ch>",
  "<http://n.pr|n.pr>",
  "<http://linkedin.com|linkedin.com>",
  "<http://uk.com|uk.com>",
  "<http://welcometothejungle.com|welcometothejungle.com>",
  "<http://gmail.com|gmail.com>",
  "<http://gmail.nl|gmail.nl>"
];

const website = nodes.var.domain
// Function to check if a website contains a generic domain using includes
{
  if (!website || typeof website !== 'string') return false;
  
  // Extract hostname from URL (remove protocol, path, etc.)
  let hostname = website.toLowerCase().replace(/^https?:\/\//, '').split('/')[0];
  
  // Check if hostname matches or is a subdomain of any generic domain
  return genericDomains.some(domain => 
    hostname === domain || hostname.endsWith('.' + domain)
  );
}

// Function to check if a website contains a generic domain using regex
{
  if (!website || typeof website !== 'string') return false;
  
  // Escape special characters in domains and join with '|'
  const escapedDomains = genericDomains.map(domain => 
    domain.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')
  );
  // Create regex: matches domain or subdomains (e.g., <http://bit.ly|bit.ly>, <http://sub.bit.ly|sub.bit.ly>)
  const regex = new RegExp(`(^|\\.)(?:${escapedDomains.join('|')})(?:$|\\/)`, 'i');
  
  // Extract hostname or test full URL
  const hostname = website.toLowerCase().replace(/^https?:\/\//, '').split('/')[0];
  return regex.test(hostname) || regex.test(website);
}
If that can helps in our fight against url shortener and co 😂 ☝️
x
Awesome, thanks for sharing! Ours is a bit shorter, will share tomorrow