Learning Continuous Hierarchies in the Lorentz Model of Hyperbolic Geometry
Abstract: We are concerned with the discovery of hierarchical relationships from large-scale unstructured similarity scores. For this purpose, we study different models of hyperbolic space and find that learning embeddings in the Lorentz model is substantially more efficient than in the Poincar\'e-ball model. We show that the proposed approach allows us to learn high-quality embeddings of large taxonomies which yield improvements over Poincar\'e embeddings, especially in low dimensions. Lastly, we apply our model to discover hierarchies in two real-world datasets: we show that an embedding in hyperbolic space can reveal important aspects of a company's organizational structure as well as reveal historical relationships between language families.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Practical Applications
Immediate Applications
Below are deployable use cases that directly leverage the paper’s findings: efficient Riemannian optimization in the Lorentz model for hyperbolic embeddings, distance-as-relatedness, norm-as-generality, and mapping to the Poincaré disk for interpretation and visualization.
- Large-scale taxonomy embedding, search, and visualization (WordNet, MeSH, ACM, EuroVoc)
- Sectors: software, healthcare, publishing, government
- What: Compress and visualize large ontologies; improve semantic search, browse-by-hierarchy, and taxonomy-aware ranking with low-dimensional hyperbolic embeddings; use norms to infer concept generality and distances to infer relatedness.
- Tools/workflows:
- Train Lorentz hyperbolic embeddings (RSGD with closed-form geodesics); map to Poincaré for UI (interactive “hyperbolic map” browsers).
- Integrate embeddings into search indices for taxonomy-aware retrieval and query expansion.
- Assumptions/dependencies: Pairwise similarity reflects same-subtree closeness; access to ontology or co-occurrence-derived similarity; basic Riemannian optimization stack (e.g., PyTorch/JAX) and negative sampling.
- Enterprise org-structure inference from communication data (email, chat, ticketing, code reviews)
- Sectors: enterprise software, HR analytics, compliance
- What: Infer latent reporting structure and role “seniority” from communication graphs; identify clusters (teams/units), and detect structural changes.
- Tools/workflows:
- Build weighted communication graphs; learn hyperbolic embeddings; monitor node norms (seniority) and cluster structures over time; dashboards for HR/ops.
- Assumptions/dependencies: Ethical and legal access to metadata-level interaction data; similarity proxies (volume, reciprocity) correlate with hierarchy; privacy-by-design and governance in deployment.
- Language family and dialect clustering from cognate or lexical similarity
- Sectors: academia (linguistics, digital humanities), education
- What: Reconstruct hierarchical relations among languages/dialects; visualize families and subfamilies; assist annotation and curriculum design for language programs.
- Tools/workflows:
- Compute normalized cognate similarities; train 2D–10D Lorentz embeddings; publish Poincaré visualizations for exploration by scholars/students.
- Assumptions/dependencies: Availability/quality of cognate or lexical similarity datasets; similarity aligns with genealogical relatedness.
- Hierarchical product catalog organization and recommendation
- Sectors: e-commerce, retail
- What: Induce or refine product hierarchies from co-view/co-purchase/title-description similarity; improve navigation, category assignment, and cold-start recommendations.
- Tools/workflows:
- Build pairwise similarities from user behavior and text embeddings; train hyperbolic embeddings; use norm to guide category depth and distance for related items.
- Assumptions/dependencies: Sufficient behavioral or content similarity; taxonomy reflects shopper mental models; periodic retraining for drift.
- Ontology alignment and taxonomy mapping (e.g., MeSH ↔ ICD, internal ↔ vendor catalogs)
- Sectors: healthcare, enterprise data management
- What: Embed multiple ontologies into a shared hyperbolic space using cross-ontology similarity to propose alignments and resolve duplicates.
- Tools/workflows:
- Construct cross-ontology similarity (string, semantic, usage); joint or aligned training; semi-automated review workflows for curators.
- Assumptions/dependencies: Comparable similarity signals across ontologies; human-in-the-loop validation for high-stakes domains.
- Hierarchical document tagging and knowledge management
- Sectors: enterprise software, publishing, legal
- What: Derive hierarchical topic structures and assign tags to documents from pairwise document similarities; enable hierarchy-aware browsing and summarization.
- Tools/workflows:
- Compute document similarities (transformers, citations); train embeddings; auto-assign tags based on nearest ancestors; visualize topic trees.
- Assumptions/dependencies: Quality document similarity; stable topic distributions; governance for taxonomy changes.
- Network and graph visualization with hyperbolic maps
- Sectors: network engineering, cybersecurity, social analytics
- What: Visualize large graphs in 2D Poincaré projections to reveal communities and central hubs; use norms to indicate node “generality/centrality.”
- Tools/workflows:
- Train Lorentz embeddings; export Poincaré coordinates; integrate into existing graph UIs (e.g., Gephi plugins, web dashboards).
- Assumptions/dependencies: Graph similarity/weight definitions reflect desired structure; embedding stability across updates.
- Hierarchy-aware classification and retrieval
- Sectors: healthcare (coding support), media, legal tech
- What: Improve multi-label classification and retrieval by penalizing errors by hierarchical distance; enable zero/few-shot via ancestor proximity.
- Tools/workflows:
- Replace or augment class embeddings with hyperbolic embeddings; train models to align sample representations to hierarchical targets; use norm priors for class generality.
- Assumptions/dependencies: Availability of a class hierarchy and consistent training data; model support for hyperbolic losses or distance-based objectives.
- Personal knowledge management and note organization
- Sectors: daily life, productivity software
- What: Automatically infer hierarchical structure in personal notes/bookmarks/photos from pairwise similarity; provide Poincaré-based “map view.”
- Tools/workflows:
- Local similarity (text/image embeddings); on-device training in small dimensions; UI to promote/demote items (adjusting positions/radii).
- Assumptions/dependencies: Usable similarity on-device; small-scale Lorentz training; user privacy controls.
- Research portfolio and scientific landscape mapping
- Sectors: academia, R&D strategy, funding agencies
- What: Map fields/subfields from citation or semantic similarity; identify central versus niche areas; support portfolio diversification decisions.
- Tools/workflows:
- Build paper/venue similarity graphs; train embeddings; monitor radial movement for generality changes; interactive landscape explorers.
- Assumptions/dependencies: Access to bibliometric/semantic data; similarity correlates with disciplinary proximity.
Long-Term Applications
These require additional research, scaling, domain adaptation, or governance frameworks before reliable deployment.
- Internet/overlay network greedy routing with hyperbolic coordinates
- Sectors: networking, edge computing
- What: Use improved low-dimensional hyperbolic embeddings for scalable, near-optimal greedy routing and topology-aware placement.
- Tools/products: “HyperRoute” modules for SDN/overlay systems; simulation-to-production pipelines.
- Assumptions/dependencies: Robust, dynamic embeddings under churn; integration with routing protocols; empirical SLAs.
- Dynamic organizational monitoring and “org design assistant”
- Sectors: enterprise software, HR/operations
- What: Track real-time changes in informal/formal hierarchies and team structures; suggest reorganizations to reduce communication bottlenecks.
- Tools/products: Org analytics platforms with continuous embedding updates and counterfactual simulations.
- Assumptions/dependencies: Streaming-safe training; strong privacy and labor-policy compliance; causal validation of recommendations.
- Clinical knowledge graphs and decision support grounded in hyperbolic hierarchies
- Sectors: healthcare
- What: Embed diseases, symptoms, procedures, and drugs to provide hierarchy-aware retrieval (e.g., differential diagnosis support), coding suggestions, and guideline navigation.
- Tools/products: CDS add-ons that surface hierarchy-consistent candidates; coding assistants informed by node norms (generality).
- Assumptions/dependencies: High-quality, curated medical ontologies and similarity metrics; rigorous clinical validation; regulatory approvals.
- Curriculum design and personalized learning pathways
- Sectors: education, edtech
- What: Represent prerequisite structures as hyperbolic hierarchies; recommend learning paths from general to specific topics tailored to learners’ profiles.
- Tools/products: LMS modules with hyperbolic maps of curricula; adaptive sequencing engines.
- Assumptions/dependencies: Accurate prerequisite/similarity estimation; pedagogical validation; fairness and accessibility considerations.
- Cross-ontology federation for interoperable data ecosystems
- Sectors: government, healthcare, enterprise data platforms
- What: Harmonize multiple evolving taxonomies/ontologies into a coherent hyperbolic space for interoperability and analytics across agencies or firms.
- Tools/products: “HyperTax Align” services with versioning, drift detection, and governance workflows.
- Assumptions/dependencies: Cross-domain similarity robustness; governance for conflicts and updates; auditability.
- Large-scale web search and ad targeting with hierarchy-aware indices
- Sectors: software, advertising
- What: Embed entities and queries into hyperbolic spaces to exploit hierarchical relations for better recall/precision and concept generalization.
- Tools/products: Hyperbolic-aware retrieval layers; query expansion using ancestor proximity; hierarchical negative sampling.
- Assumptions/dependencies: Scalability to billions of items; latency constraints; careful bias/fairness auditing.
- Phylogenetic and evolutionary modeling with hyperbolic priors
- Sectors: bioinformatics, evolutionary biology
- What: Use hyperbolic embeddings as priors or constraints in tree inference from sequence similarity; visualize and compare evolutionary trees.
- Tools/products: Hybrid phylogenetics toolkits that integrate hyperbolic embeddings with probabilistic models.
- Assumptions/dependencies: Methodological integration with likelihood-based phylogenetics; handling horizontal gene transfer; domain validation.
- Social influence and content moderation via hierarchy-aware community detection
- Sectors: social platforms, policy
- What: Identify hierarchical influence structures and community “generality” to inform moderation, recommendation throttling, or transparency reports.
- Tools/products: Platform analytics with hyperbolic embeddings; policy dashboards showing norm and distance dynamics.
- Assumptions/dependencies: Ethical frameworks; avoidance of overreach/misclassification; transparency and appeals processes.
- Software architecture inference and refactoring guidance
- Sectors: software engineering
- What: Embed modules/classes/packages from dependency and co-change similarities to infer layered architectures and refactoring targets (general core vs. specific leaf modules).
- Tools/products: IDE plugins/CI checks that surface hierarchy violations and coupling hotspots.
- Assumptions/dependencies: Reliable similarity from code graphs and histories; developer acceptance; incremental adoption.
- Macroeconomic and financial taxonomy mapping for systemic risk analysis
- Sectors: finance, policy
- What: Map instruments, institutions, and exposures into hierarchical spaces to understand contagion pathways and concentration at different “levels.”
- Tools/products: Supervisory analytics for regulators; internal risk dashboards using hyperbolic clustering.
- Assumptions/dependencies: Access to granular, timely data; sound similarity definitions; regulatory alignment and confidentiality.
Notes on feasibility across applications:
- Core dependencies: access to meaningful pairwise similarity signals; the assumption that within-subtree pairs are more similar than cross-subtree pairs; and that general items are similar to many others (captured by small norms).
- Technical considerations: Riemannian optimization capability (Lorentz model with closed-form exponential map), negative sampling for scalability, and stable mappings to Poincaré for visualization.
- Governance: privacy, ethics, and domain validation are critical in people-centric, clinical, and platform settings.
- Performance: the Lorentz-based method’s strength in low dimensions favors mobile/edge inference and interactive visualization but still requires monitoring for drift and recalibration.
Glossary
- arcosh: The inverse hyperbolic cosine function, used to express hyperbolic distances. "d_\ell(x, y) = \arcosh(-{x, y})"
- argmin: The argument (input) at which a given function attains its minimum value. "\phi(i, j) = \argmin_{k \in \mathcal{N}(i, j)} d(u_i, u_k)"
- cognates: Words in different languages that share a common origin (not via borrowing) and indicate historical relatedness. "so-called cognates, i.e., words that are shared across different languages (but not borrowed)"
- combinatorial embedding approach: A discrete, graph-structure-driven method for embedding data, emphasizing combinatorial properties. "a new combinatorial embedding approach"
- conformality: A mapping property that preserves angles locally, important in models like the Poincaré ball. "due to its conformality and convenient parameterization."
- Crowd Kernels: A method for learning similarity metrics from crowdsourced comparisons. "and Crowd Kernels"
- diffeomorphism: A smooth, invertible mapping with a smooth inverse, used to map between equivalent geometric models. "via the diffeomorphism $p : #1{H}^n \to #1{P}^n$, where"
- exponential map: A map sending a tangent vector at a point on a manifold to the manifold along the geodesic starting at that point. "The exponential map $\exp_{x} : T_{x}#1{M} \to #1{M}$"
- Gaussian Embeddings: Representations that model items (e.g., words) as Gaussian distributions to capture uncertainty and asymmetry. "proposed Gaussian Embeddings to learn improved representations."
- Generalized Non-Metric MDS: A version of multidimensional scaling that preserves rank orderings of distances without assuming metric properties. "Generalized Non-Metric MDS"
- geodesically convex cones: Cone-shaped subsets that are convex with respect to geodesics, used to model asymmetric relations in hyperbolic space. "using geodesically convex cones to model asymmetric relations."
- geodesics: The shortest (locally length-minimizing) paths on a curved manifold. "closed-form computation of the geodesics on the manifold."
- hyperbolic space: A space of constant negative curvature, well-suited for modeling hierarchical structures. "Hyperbolic space is the unique, complete, simply connected Riemannian manifold with constant negative sectional curvature."
- hyperboloid: A quadric surface; in the Lorentz model, hyperbolic space is the upper sheet of a two-sheeted hyperboloid. "denotes the upper sheet of a two-sheeted -dimensional hyperboloid"
- hypernymy: An “is-a” hierarchical relation between concepts (e.g., animal is a hypernym of dog). "provides hypernymy (is-a) relations."
- isometry: A distance-preserving transformation between metric spaces or models. "preserve all geometric properties including isometry."
- Lorentz model (of hyperbolic geometry): A model of hyperbolic space realized as a hyperboloid embedded in Minkowski space with the Lorentzian metric. "the Lorentz model of hyperbolic geometry."
- Lorentzian scalar product: An indefinite inner product with one negative and the rest positive components, defining the geometry in the Lorentz model. "denote the Lorentzian scalar product."
- Mean Average Precision (MAP): An information retrieval metric averaging precision across ranked queries. "MAP = Mean Average Precision"
- Multi-Dimensional Scaling (MDS): A technique to embed items into a low-dimensional space to preserve pairwise dissimilarities. "Multi-Dimensional Scaling (MDS) in hyperbolic space."
- ontology: A structured, often hierarchical representation of concepts and their relationships. "is a hierarchical ontology which is used by various ACM journals to organize subjects by area."
- Order Embeddings: Embeddings that encode partial order relations by enforcing coordinate-wise ordering constraints. "Another related method is Order Embeddings"
- partial order: A binary relation that is reflexive, antisymmetric, and transitive, used to formalize hierarchies. "defines a partial order over the elements of ."
- Poincaré ball model: A conformal model of hyperbolic space represented as the open unit ball with a specific metric. "the Poincaré ball model"
- Poincaré distance: The geodesic distance in the Poincaré ball model, which grows rapidly near the boundary. "numerical instabilities that arise from the Poincaré distance."
- Poincaré embeddings: Learned representations that embed hierarchical data into the Poincaré ball to exploit hyperbolic geometry. "Poincaré embeddings \citep{nickel2017poincare}"
- pseudo-Riemannian space-time: A manifold with an indefinite metric tensor (like in relativity), related to hyperbolic models used for embeddings. "pseudo-Riemannian space-time"
- Riemannian gradient: The gradient of a function defined on a manifold, computed with respect to the manifold’s metric. "denotes the Riemannian gradient"
- Riemannian manifold: A smooth manifold equipped with an inner product on each tangent space that varies smoothly. "A Riemannian manifold $(#1{M}, g)$ is a real, smooth manifold"
- Riemannian optimization: Optimization techniques that account for manifold geometry, e.g., using geodesics and manifold-aware gradients. "perform Riemannian optimization very efficiently."
- Riemannian Stochastic Gradient Descent (RSGD): A stochastic optimization method that updates parameters along geodesics using the Riemannian gradient. "Riemannian Stochastic Gradient Descent"
- sectional curvature: The curvature associated with two-dimensional sections of a manifold; negative in hyperbolic space. "constant negative sectional curvature."
- Spearman rank-order correlation: A nonparametric measure of the monotonic association between two ranked variables. "Spearman rank-order correlation "
- tangent space: The linear space of directions tangent to a manifold at a given point. "the associated \emph{tangent space}."
- transitive closure: The closure of a relation that adds edges to include all reachable pairs via path concatenation. "the undirected transitive closure of these taxonomies"