Tools/Suggestions
This page provides a collection of useful tools, scripts, and suggestions for CMS analysis workflows. Here you can find guidance for data handling, command-line utilities, and example scripts to streamline common tasks in a CMSSW or UNIX environment.
CMS UNIX (lxplus) tools
Set-up SSH with GitLab
To set up SSH access to GitLab from lxplus, generate an SSH key pair and add the public key to the GitLab account. It is useful to have a separate key specifically for CERN GitLab rather than reusing an existing SSH key. The pair of keys can be generated using the following command in the lxplus terminal.
ssh-keygen -t ed25519 -C "username@cern.ch" -f ~/.ssh/id_ed25519_cern_gitlab
Here,
id_ed25519_cern_gitlab
is the name of the private key file, and
username@cern.ch
is the email address associated with the key. The private key must be kept secure and should not be
shared. The public key will be generated in the same directory with a `.pub` extension. ED25519 is a
modern public-key signature algorithm supported by OpenSSH. It provides strong security with a compact key
and is generally preferred over generating a new RSA key. To sync this key with GitLab, the contents of
the public key file (e.g.,
~/.ssh/id_ed25519_cern_gitlab.pub
) should be copied and added to the GitLab account under
"Settings" -> "SSH Keys"
. After adding the public key, start the SSH agent and add the private key to it using the following
commands.
eval "$(ssh-agent -s)"
ssh-add ~/.ssh/id_ed25519_cern_gitlab
ssh-add -l
Then, configure the SSH client to use this key for GitLab by editing the
~/.ssh/config
file and adding the following lines.
Host gitlab.cern.ch
HostName gitlab.cern.ch
User git
Port 7999
IdentityFile ~/.ssh/id_ed25519_cern_gitlab
IdentitiesOnly yes
Then set appropriate permissions for the SSH config file.
chmod 600 ~/.ssh/config
Finally, test the SSH connection to GitLab using the following command.
ssh -T git@gitlab.cern.ch
If everything is set up correctly, a message like
Welcome to GitLab, @username!
should be displayed, confirming that the SSH connection is successful.
Finding files/datasets using DAS CLI
The CMS DAS CLI lets users quickly find datasets and files within the CMS computing system. It works on any lxplus session with a valid VOMS proxy. Instructions can be found in the CMS DAS web-page.
- Simple DAS query that displays all the available datasets using wildcards:
dasgoclient --query="dataset=/Muon*/Run2023*-22Sep2023_v*/NANOAOD" | sort -V
dasgoclient --query="file dataset=/Muon0/Run2023C-22Sep2023_v1-v1/NANOAOD"
It lists all the individual files in the dataset, with each file path on a separate line.
dasgoclient --query="file dataset=/Muon0/Run2023C-22Sep2023_v1-v1/NANOAOD" | wc -l
It counts how many files are in the dataset, and outputs a single integer representing the number of files.
dasgoclient --query="summary dataset=/Muon0/Run2023C-22Sep2023_v1-v1/NANOAOD"
It provides metadata like number of events, total size, and creation date, in key-value format.
Directly downloading files using xrdcp
Find the tutorial for the XRootD dilesystem
here
. Downloading one file is easy. Find the full path of the file using DAS CLI commands and run
xrdcp
as follows.
xrdcp root://cmsxrootd.fnal.gov//store/data/Run2023C/Muon0/NANOAOD/22Sep2023_v1-v1/2530000/030eead5-93f9-405c-863c-a62244712e91.root .
I find it helpful to define the following function in my
~/.bashrc
file.
get_from_das() {
[ -z "$1" ] && echo "Usage: get_from_das <filename> [outputname]" && return 1
cmd="xrdcp root://cmsxrootd.fnal.gov/$1 ${2:-.}"
echo -e "\033[93mCommand:\033[0m \033[3;33m$cmd\033[0m"
eval "$cmd"
}
Usage:
get_from_das /store/data/Run2023C/Muon0/NANOAOD/22Sep2023_v1-v1/2530000/030eead5-93f9-405c-863c-a62244712e91.root
Here I am using the redirector
cmsxrootd.fnal.gov
(server that acts as a middleman between client and the actual storage site) which is based in Fermilab.
The following alternative redirectors are also available.
-
cmsxrootd-global.cern.ch(global CMS XRootD redirector at CERN) -
xrootd-cms.infn.it(INFN, Italy)
For downloading files in bulk, I made a setup which can be found in GitHub: xrdcp-DAS-to-EOS . This setup needs the DAS names as inputs, uses DAS CLI to find the files and xrdcp to download them. Be careful before downloading large datasets, as it may take a long time and consume significant bandwidth. This setup is helpful for downloading small MC samples quickly, avoiding any CRAB-job submission.
Luminosity calculation using run-number and luminosity-section (JSON)
The
brilcalc
tool is available in lxplus at
~/.local/bin/brilcalc
. It takes a JSON file as input, which contains run-number and lumisections in the following format.
{
"321975": [[591, 593], [597, 599], [717, 734] , [736]],
....
}
This tool is run as follows.
brilcalc lumi \
--normtag /cvmfs/cms-bril.cern.ch/cms-lumi-pog/Normtags/normtag_PHYSICS.json \ # Path to the normtag JSON with luminosity calibration constants
-c web \ # Use web API as input source (avoids old Oracle DB issues)
-u /fb \ # Output integrated luminosity in femtobarns (/fb)
-i [[json file name]] # Input golden JSON file with certified lumisections
While running brilcalc, it is recommended (by LUM POG) to use a physics approved normtag (in this case, normtag_PHYSICS.json). This normtag json converts the numbers into physical units.
- More information on from LUM POG: LUM POG TWiki page
- Golden JSONs: cms-service-dqmdc.web.cern.ch/CAF/certification/
- GitHub repository with summary of luminosities and era-wise Golden JSONs: luminosity-calculation-per-era
Cross-section calculation using GenXSecAnalyzer
The GenXSecAnalyzer is used to calculate or extract the cross-section of a simulated process from the LHE level. I made a wrapper script to run it for a list of samples. The detailed instructions are in this documentation . The tool described in the official documentation requires MiniAOD files as inputs. My wrapper simply requires DAS names of the target NanoAOD samples, along with a CMSSW release and CMS VOMS proxy.
cmsrel CMSSW_13_0_13
cd CMSSW_13_0_13/src
cmsenv
voms-proxy-init -voms cms
After setting up, bring the
find_xsec_fromDAS.py
file here. It takes a DAS name of the NanoAOD sample as input.
python3 find_xsec_fromDAS.py --dataset [DAS name of the NanoAOD sample]
-
Brings
genXsec_cfg.pyinto the work area (if not already present).curl https://raw.githubusercontent.com/cms-sw/genproductions/master/Utilities/calculateXSectionAndFilterEfficiency/genXsec_cfg.py - Uses DAS CLI to locate the parent (MiniAOD) dataset.
- Selects a MiniAOD file.
-
Runs the cross-section analyzer:
cmsRun genXsec_cfg.py inputFiles=file:root://cms-xrd-global.cern.ch//[filepath] - Extracts the relevant value from the technical jargon and displays it on screen.
Submitting CRAB jobs and bringing samples to EOS
The CRAB framework manages distributed computing tasks by packaging user-defined analysis codes with configuration files that specify input datasets and output locations. Upon submission, jobs are distributed across the WLCG using a scheduler for dynamic resource allocation. It monitors job status through a database, allowing users to track progress and retrieve output files and logs, ensuring integration with CMS data management for efficient data analysis. Detailed instructions can be found here . I inherited the CRAB setup from Arnab , and tweak it around a bit. The setup is currently available in GitHub:phazarik/MakeSelector-CRAB-setup . I use it to skim nanoAOD files using MakeSelector based codes and dump them in my EOS space before bringing them to a local device. Go to my setup and follow these steps.
- Login to lxplus8.
- Be in a CMSSW environment. In my case, its CMSSW_13_0_13, but the exact version is not crucial. This is just needed for loading some utilities for submitting CRAB jobs.
- Generate voms-proxy for CMS to access the files from DAS.
-
Clone the following repository in your work area.
git clone https://github.com/phazarik/MakeSelector-CRAB-setup -
Run
./compile_and_run.Clocally by feeding it some local NanoAOD file to test the MakeSelector based source code. -
Submit one job:
If previous step runs smoothly and produces an output file, you can test-submit one job by running
crab_config.py. This submits a crab-job for one dataset and manages howcrab_script.shshould run remotely. The output directory path have to be changed in the script as needed. Some parameters need to be changed externally. The following is an example of how to run the setup for one job.export CRAB_CAMPAIGN=2016preVFP_UL export CRAB_SAMPLENAME=Muon_B1 export CRAB_FLAG=muon crab submit crab_config.py \ # submit using the CRAB configuration file General.requestName=nanoSkim_2016pre_Muon_B1 \ # unique name for this CRAB task Data.inputDataset=/SingleMuon/Run2016B-ver1_HIPM_UL2016_MiniAODv2_NanoAODv9-v2/NANOAOD # DAS name of the input dataset (NanoAOD) -
Bulk submission:
I keep dataset names, FLAG and SAMPLENAME in external text files in
samplelistsdirectory. parameters externally. Then I use thebulkSubmitCrab.pyscript to read those files and submit multiple jobs in one go. Edit this script as needed to select the desired campaign and list of samples. The usage is as follows.python3 bulkSubmitCrab.py --dryrun False
-
crab submitsendscrab_config.py,crab_script.sh, and analysis code to the grid. -
On the worker node,
crab_script.shsets up CMSSW and prepares the environment. -
Input file assignment is read from
PSet.py. -
crab_script.shinvokesshell_instructions.shwith campaign, sample, and flag arguments. -
shell_instructions.shrunscompile_and_run.C, producingskimFile.root. -
CRAB transfers
skimFile.rootand logs to the output location incrab_config.py.
NanoAOD based tools
Investigating NanoAOD file structure
I made this tool to checkout and compare the branch types of different NanoAOD formats. For a given input file of a particular NanoAOD structure, it finds out the "Events" branch which is relevant to us. Then looping over each branch, it accesses the name and leaves associated with the branch and prints out the type. The code can be found here . Usage:
[root].x find_nanoAOD_structure("filename.root");
Generating an analysis template
For a given NanoAOD file, a
MakeSelector()
based template can be generated in the following way. Make sure to note down the NanoAOD version, because
this template may have some branch-mismatch issue with other versions. Read the ROOT file and do the
following in the ROOT prompt.
[root] TFile *f = new TFile("filename.root");
[root] gROOT -> FindObject("Events");
[root] Events -> MakeSelector("anaName"); #Pick an analysis name
This should produce a template analyzer class named
anaName
. The class is kept in a header file, and its functions including the event loop is run in a C file. This
same approach works for any custom TTree containing collision events.
Histogram-maker
The following is a standalone analysis tool for processing nanoAOD files. The MakeSelector-based source code generates histograms for physics variables, and the utility tools allow overlaying multiple histograms for comparison. This is tested on nanoAOD versions: v7, v9 (RunII-UltraLegacy) and v12 (Run3Summer22). The header files allow switching between different branch types, which accounts for the branch type changes from NanoAODv11 to v12. The number of branches being read is kept minimal to avoid conflicts between different versions. Run-instructions are available on the Git page.
GitHub: https://github.com/phazarik/nanoAOD_analyzer
TIFR T2/T3 tools
Accessing and copying files to T3 [in bulk] using xrdcp
The T2 filesystem does not allow me to do
ls
easily and wildcards are not allowed. I also can't use python features like
os.listdir()
. I use python scripts to run commands such as
xrdfs se01.indiacms.res.in ls
(or
xrdcp
) iteratively from the T3 area and list/copy all the root files in a given directory.
-
findPathsT2.py: Finds full paths to the root files in a given directory. Usage:python3 findPathsT2.py /store/user/alaha/nanoRDFjobs -
getFilesT2.py: Copies all the root files from a given T2 directory to the T3 area. Usage:python3 getFilesT2.py base_directory_in_T2 output_directory_in_T3 -
getIndividualFilesT2.py.txt: In case the above method fails, or only some specific files are needed, this script brings them by reading a text file. It manually constructs the full path as given in the text file instead of using xrdfs ls. Usage:
where, the text file has the following format.python3 getIndividualFilesT2.py.txt output_directory_in_T3/store/user/alaha/nanoRDFjobs/SingleMuon/NanoRDF__20240916_095804/240916_075807/0000/ntuple_skim_3.root /store/user/alaha/nanoRDFjobs/Muon/NanoRDF__20240916_095809/240916_075812/0000/ntuple_skim_102.root /store/user/alaha/nanoRDFjobs/Muon/NanoRDF__20240916_095809/240916_075812/0000/ntuple_skim_109.root /store/user/alaha/nanoRDFjobs/Muon/NanoRDF__20240916_095809/240916_075812/0000/ntuple_skim_119.root /store/user/alaha/nanoRDFjobs/Muon/NanoRDF__20240916_095815/240916_075817/0000/ntuple_skim_13.root
Deleting files from T2
Deleting stuff from T2 is tricky.
rm
only works for individual files.
rmdir
only works for empty directories. I came up with the following method. The python file takes a list of
directories in this format:
/store/user/username/directory
.
- Listing all the folders, subfolders and files.
- Dividing the list into two: one for files, one for directories.
- Deleting the individual files first, by looping over the list of files.
- Deleting the directories, starting from the deepest ones.
For this, use cleanT2.py and mention the directories you want to remove in the list inside.
VOMS proxy hack
As of August 2024, the cms-proxy fails to generate at the T3 area. That's why I am manually creating the proxy file in lxplus, and bringing it to ui.indiacms.res.in. Before bringing the file to T3, the file has to be transferred to an accessible ares in lxplus.
-
Create the proxy in lxplus area:
voms-proxy-init -voms cms -valid 192:00 voms-proxy-info -all cp /tmp/x509up_u139657 . # Give the correct proxy filename -
Bring the proxy into T3 area:
scp phazarik@lxplus.cern.ch:~/x509up_u139657 . # Give correct username realpath x509up_u139657 # copy this export X509_USER_PROXY=/grid_mnt/t3home/phazarik/x509up_u139657 # Give the copied path here
Miscellaneous tools
Making Feynman diagrams in LaTeX
When it comes to drawing Feynman diagrams in LaTeX using feynmf and TikZ-Feynman, the later is generally
easier and more flexible to use. The first one is older and requires a more complicated setup, while
TikZ-Feynman works smoothly with modern LaTeX with more customization options. If you're just starting
out, TikZ-Feynman is the better choice. A template for drawing Feynman diagrams using tikz-feynman package
can be found
here
. It already has some examples. It can be compiled using
pdflatex
as follows.
pdflatex feynman-template.tex
Just make sure that you have the necessary LaTeX packages installed beforehand. I prefer to produce individual pdf files for each Feynman diagram and convert them into high quality png files. The following is an example of a VLL production diagram.
\documentclass[tikz,border=3mm]{standalone}
\usepackage{tikz-feynman}
\tikzset{every picture/.style={line width=1.1pt}}
\begin{document}
\begin{tikzpicture}[baseline={(current bounding box.center)}]
\begin{feynman}
\vertex (v1);
\vertex [right =1.5cm of v1] (v2);
%incoming vertices
\vertex [above left =1.5cm of v1] (i1) {\(q\)};
\vertex [below left =1.5cm of v1] (i2) {\(\bar{q}\)};
%internal vertices connected to the v2
\vertex at($(v2)+(1.2, +0.7)$) (b1);
\vertex at($(v2)+(1.2, -0.7)$)(b2);
%internal vertices connected to b1 and b2
\vertex at($(b1)+(1.2, +0.1)$) (c1);
\vertex at($(b2)+(1.2, -0.1)$) (c2);
%outgoing vertices
\vertex at($(b1) + (1.0, +1.2)$) (o1) {\( l \)};
\vertex [above right =0.7cm of c1] (o2);
\vertex [below right =0.7cm of c1] (o3);
\vertex [above right =0.7cm of c2] (o4);
\vertex [below right =0.7cm of c2] (o5);
\vertex at($(b2) + (1.0, -1.2)$) (o6) {\( \nu \)};
\diagram*{
%incoming lines
(i1) -- (v1) -- (i2);
%internal lines
(v1) --[boson, color=black, edge label = \({\color{black} Z/\gamma^*}\)] (v2);
(b1) -- (v2) -- (b2);
(b1) --[boson, edge label' = \(Z\)] (c1);
(b2) --[boson, edge label = \(W\)] (c2);
%outgoing lines
(o1) -- (b1);
(o2) -- (c1) -- (o3);
(o4) -- (c2) -- (o5);
(b2) -- (o6);
};
%labels (manually putting them here because too crowded)
\vertex at($(o2) + (+0.3, +0.0)$) (l2) {\(q^{\prime}\)};
\vertex at($(o3) + (+0.3, +0.0)$) (l3) {\(\bar{q}^\prime\)};
\vertex at($(o4) + (+0.3, +0.0)$) (l4) {\( q^{\prime\prime}\)};
\vertex at($(o5) + (+0.3, -0.0)$) (l5) {\( \bar{q}^{\prime\prime}\)};
\vertex at($(b1) + (-0.6, +0.0)$) (tau1) {\( E \)};
\vertex at($(b2) + (-0.6, -0.0)$) (tau2) {\( E \)};
\end{feynman}
\end{tikzpicture}
\end{document}
This example only requires the tikz-feynman package. The pdf output can be converted to png as follows.
# Note: Don't use the -alpha remove option if you want transparent png files.
convert -density 300 vll_production.pdf \
-colorspace RGB -quality 90\
-alpha remove -background white\
vll_production.png