Changelog
Source:NEWS.md
realestatebr 1.2.0
CRAN release: 2026-10-01
- Added
query_dataset("mcmv"), with Minha Casa, Minha Vida financing contracts, official financing totals, and federally subsidized housing projects from the Ministério das Cidades. A new article covers the three tables, and a data dictionary documents their columns and source changes. - Added
get_dataset("paic"), with annual IBGE construction-industry data for the new series starting in 2024 (SIDRA tables 10463, 10441, and 10442). Theactivity,size, andstatetables cover enterprises, employment, revenue, costs, and output; values publish in original units with avalue_statuscolumn for SIDRA symbols. Do not compare with the 2007-2023 series (see IBGE Nota técnica 01/2026). - Added
get_dataset("pnad_housing")with annual PNAD Contínua household tenure, dwelling type, household size, mean size, and domestic-unit composition estimates, including published shares and coefficients of variation in their original units. - Fixed PNAD Housing validation to reject comma decimals that the numeric parser cannot read, and updated live coverage tests to accept new published years.
- Fixed the cache publication workflow to save its targets checkpoint after successful publication, allowing failed uploads to retry on the next run.
realestatebr 1.1.0
- Added
get_dataset("pim_pf_construction"), a linked monthly IBGE production index for construction inputs from January 1991. - Added
get_dataset("sinapi"), with monthly IBGE construction costs and indices by state, region, and Brazil, both with and without payroll-tax relief. - Added
query_dataset()for lazy, joinable DuckDB access to large relational datasets. Its first dataset iscno, with annual snapshots of the Federal Revenue Service construction registry published as Parquet files on GitHub Releases.
Bug fixes
- Unknown-dataset errors from
get_dataset()now list only materialized datasets in alphabetical order. -
get_dataset("secovi", source = "fresh")andget_dataset("bcb_realestate", source = "fresh")now raise an error when the original source is unavailable instead of silently returning the GitHub release cache. -
get_dataset("secovi")now downloads each indicator page separately with a timeout and user agent, and warns about indicators it could not read. - Download retry warnings and errors now show the underlying cause instead of a generic “In index: 1.” message.
- The weekly cache pipeline keeps SECOVI series from the published cache when they are missing from a fresh download, and reports target warnings as workflow annotations.
-
get_dataset("rppi", source = "github")now supports the combinedalltable and every individual RPPI table through the dataset release cache. - The weekly cache pipeline now refreshes age-cued datasets reliably despite schedule jitter, uploads only cache files produced from updated upstream targets, rejects stale SECOVI data, and reports target failures as failed workflow runs.
-
get_dataset("secovi")once again downloads current SECOVI-SP data afterxml21.6.0 changed its default HTML encoding. SECOVI responses are now parsed explicitly as ISO-8859-1, and upstream indicators 25 (launches_rmsp) and 118 (sales_rmsp) are no longer requested because SECOVI has discontinued them.
Other changes
- Replaced the
rbcbdependency withGetBCBData, which splits long requests into windows the BCB SGS API accepts. -
get_dataset("bcb_series")now returns each series in full, from its first published observation instead of 2010-01-01. Daily series such as SELIC (code 432) were previously dropped, since SGS rejects requests spanning more than ten years. - Added
quiettoget_dataset(). -
get_dataset()now reports the dataset, the table, and the source it came from in a single message. Setoptions(realestatebr.debug = TRUE)for step-by-step progress. - Fixed
get_dataset("bcb_series")failing withobject 'bcb_metadata' not foundwhen the package was not attached (#23). The IVAR series inget_dataset("rppi")silently fell back to the GitHub release for the same reason.
realestatebr 1.0.1
CRAN release: 2026-06-05
CRAN resubmission. Version 1.0.0 was archived on 2026-05-28 because the package and its transitive piggyback/gh dependencies wrote to ~/.cache/realestatebr/ and ~/.cache/R/gh/, violating CRAN’s policy on home-filespace writes. This release removes this: the caching architecture is fully overhauled.
Caching architecture
The user-level disk cache has been removed. The package no longer writes to ~/.cache/realestatebr/ (or any other location outside the R session’s temporary directory).
-
get_dataset()is now two-tier: it tries the package’s GitHub release asset first and falls back to a fresh download from the original source. Thesourceargument no longer accepts"cache"; valid values are"auto","github", and"fresh". Themax_ageargument has been removed (cache freshness is now managed by the weekly CI/CD release pipeline). - Repeated
get_dataset()calls within one R session are served from a package-private in-memory environment. Use the newclear_session_cache()function to drop it. - The exported helpers
clear_user_cache(),check_cache_status(), andupdate_cache_from_github()have been removed. - Internal dataset functions (
get_abecip_indicators(),get_secovi(),get_rppi_*(), etc.) no longer accept acachedargument. They always download from the original source; useget_dataset(source = "github")for the pre-processed asset. -
piggyback(formerly Suggests) andrappdirs(formerly Imports) have been dropped. GitHub release assets are now fetched directly viahttr::GET()against the public release-asset URL, avoiding the transitivegh::gh()cache writes that also violated CRAN policy.
Documentation
- Every dataset now has its own help topic (
?abecip,?bcb_realestate,?rppi, and so on) documenting available tables, columns, units, and coverage. The topics are generated from the dataset registry (inst/extdata/datasets.yaml), which now carries column-level metadata, viadata-raw/generate_dataset_docs.R. - Fixed stale
dim_citydocumentation: the table has 9 columns, including the previously undocumentedabbrev_stateandyear. - Two new website-only articles, “Housing Credit in Brazil” (
abecip,bcb_series,bcb_realestate) and “The Primary Market and the Construction Cycle” (abrainc,secovi,fgv_ibre).
Bug fixes
-
attach_dataset_metadata()now accepts"bundled"as a validsourcevalue, fixing a hard error whenget_dataset("abecip", "cgi")was called withsource = "fresh". The CGI table loads from a bundledinst/extdatafile and was previously tagged with the rejected value"cache".
realestatebr 1.0.0
CRAN release: 2026-05-27
Breaking changes
-
nre_irehas been removed. It required fully manual updates and had no automated pipeline support; it will be reconsidered for a future release. -
property_recordshas been removed because the upstream data source is no longer available. -
cbichas been removed. The upstream CBIC portal migrated to a restricted-access platform. The five cement tables will be rebuilt from IBGE open data in a future release. -
itbi_summaryand the internal ITBI helpers (get_itbi,get_itbi_bhe) have been removed. They were incomplete (single-municipality coverage) and are deferred to a future version.
Changes to bcb_series
-
get_dataset("bcb_series")now returns only four columns:date,code_bcb,name_simplified, andvalue. Full metadata is available viabcb_metadata. - The
tableargument now accepts a hierarchy level:"core"(default),"primary","secondary","tertiary", or"full". The levels are cumulative —"primary"includes all"core"series plus key macro indicators such as SELIC, IPCA, and INCC. Previously the argument accepted BCB category names ("credit","price", etc.). -
bcb_metadatagains ahierarchycolumn (integer 1–4) that records the relevance tier assigned to each series.
Changes to rppi
-
get_dataset("rppi", table = "all")now returns two additional columns:transaction_type("sale"or"rent") andsource(the index name, e.g.,"IGMI-R","IVG-R","FipeZap"). Previously the stacked table had no way to distinguish transaction type from index source. - Fixed a sheet-encoding mismatch in the FipeZAP XLSX download that caused
Sheets not founderrors after FIPE added a new summary sheet. Sheet selection now uses numeric indices to avoid Latin-1/UTF-8 mismatches in accented sheet names. - Fixed
get_range()failing with a malformed cell range whentidyxlreceived a sheet name with a mismatched encoding attribute. The function now derives the sheet lookup key directly from the cells returned bytidyxlrather than from the user-supplied string.
Bug fixes
- Restored
fetch_fgv_local(), which was inadvertently dropped during earlier refactoring but is still called by thetargetspipeline to process the manually-maintained FGV IBRE CSV export.
Internal changes
- Dataset download functions now follow a
download_*/clean_*naming convention:download_*functions return a file path;clean_*functions parse and tidy that path into a tibble. - Replaced bare
tryCatchwithrlang::try_fetchthroughout, usingparent = cndto preserve the original error chain. - All dataset functions now call
validate_dataset_params(),attach_dataset_metadata(), andvalidate_dataset()fromR/helpers_dataset.RandR/helpers_download.Rinstead of re-implementing these patterns inline.
CI / pipeline
- Fixed stale target names in
.github/workflows/update_data_weekly.yml. The workflow was silently skippingabecipandabrainctargets on every automated run; those entries now use the current granular target names (abecip_sbpe_data,abecip_units_data,abrainc_indicator_data, etc.). - Added
fgv_ibre_fileto the weekly andalltarget groups so the FGV file-change target is included in scheduled runs.
Documentation
- Updated the
bcb_seriestable reference invignettes/getting-started.Rmdto use the new hierarchy levels ("core","primary","secondary","tertiary","full") instead of the removed category names. - Updated the
rppi_bistable listing to includedetailed_annualanddetailed_halfyearly, which were previously omitted.
realestatebr 0.6.1
CRAN Submission Fixes
- Upgraded HTTP URLs to HTTPS in
get_secovi.R - Added
URLfield to DESCRIPTION - Fixed incomplete
@sourcetag fordim_citydataset documentation - Fixed typo in
b3_real_estatedocumentation (“mian” -> “main”) - Removed local
skip_on_cran()definition that shadowedtestthat - Updated stale description for CBIC dataset in registry
realestatebr 0.6.0 (2025-11-09)
Cache Freshness Detection (Phase 2)
Version 0.6.0 introduces an intelligent cache freshness detection system with relaxed defaults to avoid annoying users with unnecessary warnings.
New Features
Cache Age Tracking
-
get_cache_age(): Returns cache age in days for any dataset -
is_cache_stale(): Checks if cache exceeds recommended freshness thresholds -
check_cache_status(): Diagnostic function showing status of all cached datasets
Relaxed Warning Thresholds
Cache warnings only appear when data is significantly stale (exceeds 2x the update frequency): - Weekly datasets: warn after 14 days (not 7) - Monthly datasets: warn after 60 days (not 30) - Manual datasets: never warn
Advanced Cache Control
-
max_ageparameter inget_dataset(): Force fresh download if cache exceeds specified age - Useful for users who need very recent data
- Most users won’t need this - relaxed defaults handle typical use cases
Examples
# Check status of all cached datasets
check_cache_status()
# Get age of specific dataset
get_cache_age("bcb_series")
# Check if dataset is stale (uses relaxed defaults from registry)
is_cache_stale("bcb_series")
# Advanced: Force very fresh data (< 1 day old)
get_dataset("bcb_series", max_age = 1)
# Advanced: Only use cache if less than 3 days old
get_dataset("rppi", table = "sale", max_age = 3)Code Simplification: Logic Consolidation (Phase 3)
Generic Helper Functions
Version 0.6.0 introduces 7 generic helper functions that consolidate 890 lines of repetitive code patterns across dataset functions.
What Changed
- Created: R/helpers-dataset.R with 6 new helper functions (430 lines)
- Refactored: 5 core files using new helpers (417 lines removed, 15.4% reduction)
- Added: apply_table_filtering() in R/get-dataset.R (95 lines, eliminates 156 lines of duplication)
- Added: 52 comprehensive tests for all helper functions
- Improved: Consistent error messages and metadata across all datasets
File-by-File Results
| File | Before | After | Lines Saved | % Reduction |
|---|---|---|---|---|
| get_abecip_indicators.R | 551 | 431 | 120 | 21.8% |
| get_abrainc_indicators.R | 544 | 445 | 99 | 18.2% |
| get_secovi.R | 438 | 356 | 82 | 18.7% |
| get_bcb_series.R | 334 | 278 | 56 | 16.8% |
| get-dataset.R | 833 | 773 | 60 | 7.2% |
| TOTAL | 2,700 | 2,283 | 417 | 15.4% |
New Helper Functions
-
validate_dataset_params() (R/helpers-dataset.R)
- Consolidated input validation for table, cached, quiet, max_retries parameters
- Ensures consistent error messages across all datasets
- Saves ~28 lines per file
-
handle_dataset_cache() (R/helpers-dataset.R)
- Unified cache loading with fallback strategies
- Consistent error handling and user messages
- Saves ~35-50 lines per file
-
attach_dataset_metadata() (R/helpers-dataset.R)
- Standardized metadata attachment (source, download_time, download_info)
- Flexible extra_info parameter for dataset-specific metadata
- Saves ~8-16 lines per file
-
validate_dataset() (R/helpers-dataset.R)
- Generic data validation (rows, columns, dates)
- Configurable validation rules with detailed error messages
- Saves ~44 lines per file
-
validate_excel_file() (R/helpers-dataset.R)
- Excel file validation (size, expected sheets)
- Used by abrainc and abecip functions
- Prevents silent failures
-
download_with_retry() (R/rppi-helpers.R - REUSED)
- Found existing implementation, avoided duplication
- Saves ~46-86 lines per file that would have been duplicated
-
apply_table_filtering() (R/get-dataset.R)
- Centralizes all table/category filtering logic
- Supports property_records, SECOVI, BCB Real Estate, BCB Series
- Eliminates 156 lines of duplication between cache functions
Files Changed
- New: R/helpers-dataset.R (430 lines, 6 helpers, 52 tests)
- Updated: R/get_abecip_indicators.R (21.8% reduction)
- Updated: R/get_abrainc_indicators.R (18.2% reduction)
- Updated: R/get_secovi.R (18.7% reduction)
- Updated: R/get_bcb_series.R (16.8% reduction)
- Updated: R/get-dataset.R (7.2% reduction, added apply_table_filtering())
BREAKING CHANGES: API Simplification (Phase 2)
Removed Deprecated Function Exports
Version 0.6.0 removes 8 deprecated functions from the public API. These functions are now internal-only.
What Changed
Removed from NAMESPACE: 8 deprecated functions no longer exported: - get_abecip_indicators() - get_abrainc_indicators() - get_bcb_realestate() - get_bcb_series() - get_fgv_ibre() - get_nre_ire() - get_rppi_bis() - get_secovi()
Impact
- Users must migrate: These functions can no longer be called directly
-
get_dataset() is the only supported API: All data access must go through
get_dataset() - Cleaner public interface: Package exports only essential user-facing functions
Migration
These functions were deprecated in v0.4.0 (18+ months ago). Users must now use get_dataset():
# Old way (NO LONGER WORKS):
data <- get_secovi()
data <- get_bcb_series(table = "price")
data <- get_abecip_indicators(table = "sbpe")
# New way (REQUIRED):
data <- get_dataset("secovi")
data <- get_dataset("bcb_series", "price")
data <- get_dataset("abecip", "sbpe")Rationale
-
Simpler API: One function (
get_dataset()) instead of 15+ - Pre-1.0.0 flexibility: Breaking changes acceptable before stable release
Code Clarity Improvements
Renamed Confusing “Legacy” Terminology
-
Renamed:
get_from_legacy_function()→get_from_internal_function() - Rationale: These functions call internal worker functions, not “legacy” code
- Impact: Internal only - no user-facing changes
Files changed: R/get-dataset.R
CBIC Code Simplification and Table Availability
Code Reduction (~223 lines, 11% smaller)
- Removed: 100+ lines of commented-out old implementation code
-
Removed: 4 unused helper functions (124 lines total):
-
suppress_external_warnings()- Never called -
explore_cbic_structure()- Only in examples -
get_cbic_files()- Only in examples -
get_cbic_materials()- Only in examples
-
-
Removed: Unnecessary metadata attributes from
get_cbic_steel()andget_cbic_pim():attr(result, "source")attr(result, "download_time")attr(result, "download_info")- Associated tryCatch blocks and cli_user messages (~69 lines)
Table Availability Fixed
-
Unblocked:
steel_pricesandpimtables now accessible -
Blocked: Only
steel_productionremains blocked (has data quality issues) - Updated: Error messages now accurately reflect v0.6.0 status
-
Available tables:
- Cement: cement_monthly_consumption, cement_annual_consumption, cement_production_exports, cement_monthly_production, cement_cub_prices
- Steel: steel_prices
- PIM: pim, pim_production_index
BREAKING CHANGES: Documentation Simplification (Phase 1)
Removed Examples from Deprecated Functions
Version 0.6.0 removes usage examples from deprecated legacy functions to simplify the codebase. Since we are pre-1.0.0, this is an acceptable breaking change.
What Changed
-
Removed: All
@examplesblocks from 8 deprecated functions: -
Removed: Verbose
@sectionblocks (Progress Reporting, Error Handling) -
Simplified:
@detailssections to 1-3 essential lines -
Enhanced:
@section Deprecationblocks with code migration examples
Impact
- ~260-290 lines of documentation removed
- Documentation now focuses on migration guidance rather than usage examples
- All functions still exported and callable (no functionality changes)
- Help pages now emphasize using
get_dataset()instead
Migration
These functions were deprecated in v0.4.0. Users should migrate to the modern API:
# Old way (still works, but no longer documented with examples):
data <- get_secovi()
# New way (recommended):
data <- get_dataset("secovi")Full migration examples are available in each function’s @section Deprecation block.
Rationale
- Pre-1.0.0: Breaking changes are acceptable before stable release
- Codebase simplification: Reduces maintenance burden and package size
-
Focus on modern API: Encourages users to adopt
get_dataset()interface - Clear migration path: Enhanced deprecation warnings guide users to new API
realestatebr 0.5.1
Bug Fixes
SECOVI Default Table Fix
Fixed SECOVI dataset to return all categories by default instead of only “condo”
Problem:
get_dataset("secovi")was only returning the “condo” category (1,939 rows) instead of all categories (9,398 rows). This caused test failures for launch/rent/sale tables.Root Cause: When no table parameter was specified, the code defaulted to the first category alphabetically (“condo”), rather than fetching all categories.
-
Solution:
- Added
default_tableconfiguration support indatasets.yaml - Updated
validate_and_resolve_table()to check fordefault_tablesetting - Set SECOVI’s
default_table: "all"in registry - Regenerated cache with all 4 categories
- Added
-
Impact:
- Cache size: 12KB → 55KB (includes all categories)
- Data completeness: 1,939 → 9,398 rows
- Categories: condo (1,939), launch (780), rent (2,779), sale (3,900)
# Now returns all categories by default
get_dataset("secovi") # → 9,398 rows, 4 categories
# Specific tables still work correctly
get_dataset("secovi", "launch") # → 780 rows
get_dataset("secovi", "rent") # → 2,779 rows
get_dataset("secovi", "sale") # → 3,900 rowsTest Infrastructure Improvements
- Updated test suite to use
devtools::load_all()instead oflibrary()to ensure testing of development version - Added comprehensive pre-release test suite (
tests/comprehensive_check_v0.5.qmd) - Added test result documentation (
tests/TEST_RESULTS_SUMMARY.md,tests/QUICK_SUMMARY.md)
realestatebr 0.5.0
BREAKING CHANGES: User-Level Caching Architecture
Major Architectural Change
Version 0.5.0 introduces user-level caching, removing bundled datasets from the package to comply with CRAN’s 5MB size limit. This is a BREAKING CHANGE that affects how datasets are accessed.
What Changed
-
Removed: All cached datasets from
inst/cached_data/(previously ~25MB) -
Added: User-level cache directory at
~/.local/share/realestatebr/(Linux/Mac) or%LOCALAPPDATA%/realestatebr/Cache/(Windows) - Added: GitHub Releases integration for pre-processed datasets
-
Changed:
source="cache"now refers to user cache, not package cache -
Changed:
source="github"now downloads from GitHub releases, not package files
New Cache Behavior
# First use: downloads from GitHub releases to user cache
data <- get_dataset("abecip") # Downloads once
# Subsequent uses: loads from user cache (instant, offline)
data <- get_dataset("abecip") # Loads from ~/.local/share/realestatebr/
# Force fresh download from original source
data <- get_dataset("abecip", source = "fresh") # Downloads and caches
# Explicit source selection
data <- get_dataset("abecip", source = "cache") # User cache only
data <- get_dataset("abecip", source = "github") # GitHub releases onlyAuto Fallback Strategy (source = “auto”, default)
-
User Cache: Check
~/.local/share/realestatebr/(instant, offline) -
GitHub Releases: Download pre-processed data (requires
piggybackpackage) - Fresh Download: Download from original source (saves to user cache)
New Dependencies
-
Added:
rappdirs(Imports) - Cross-platform user cache directory support -
Added:
piggyback(Suggests) - GitHub releases download support
New Functions
-
get_user_cache_dir(): Get path to user cache directory -
list_cached_files(): List all cached datasets -
clear_user_cache(): Remove cached datasets -
is_cached(): Check if dataset is in cache -
list_github_assets(): List available datasets on GitHub releases -
download_from_github_release(): Download specific dataset from releases -
update_cache_from_github(): Update cached datasets from GitHub -
is_cache_up_to_date(): Compare local vs GitHub cache timestamps
Migration Guide
For Users
# Install updated package
install.packages("realestatebr") # or devtools::install_github()
# Install piggyback for GitHub downloads (recommended)
install.packages("piggyback")
# First use after update: will download datasets to user cache
data <- get_dataset("abecip")
# Check cache location
get_user_cache_dir()
# Manage cache
list_cached_files() # See what's cached
clear_user_cache("abecip") # Clear specific dataset
clear_user_cache() # Clear all (with confirmation)For Package Developers
- Cached data files now excluded from package build via
.Rbuildignore - Package size reduced from ~25MB to <5MB (CRAN compliant)
-
inst/cached_data/kept for development/CI but excluded from distribution - GitHub Actions workflow publishes cache to releases via
data-raw/publish-cache.R
Deprecations
-
import_cached(): Still works but now loads from user cache (previously frominst/) - Old
cached=TRUEparameter in legacy functions: Still supported but uses new cache
Files Changed
-
New:
R/cache-user.R- User cache management -
New:
R/cache-github.R- GitHub releases integration -
New:
data-raw/publish-cache.R- Upload cache to releases -
Updated:
R/get-dataset.R- Refactored cache logic -
Updated:
R/cache.R- Marked as deprecated (kept for compatibility) -
Updated:
.Rbuildignore- Excludeinst/cached_data/files -
Updated:
DESCRIPTION- Addedrappdirsandpiggybackdependencies
Targets Pipeline Fixes
Critical Pipeline Functionality
- Fixed: Targets pipeline now fully functional for automated data updates
-
Fixed: FGV IBRE and NRE-IRE datasets now work correctly in targets pipeline
- Changed from
source="fresh"tosource="github"for manually-updated datasets - These datasets have no API/download capability and require manual updates
- Changed from
-
Fixed: Removed broken internal data object fallback in
get_fgv_ibre()andget_nre_ire()- Previously tried to access non-existent
fgv_dataandireobjects fromR/sysdata.rda - Now provides clear error messages when fresh downloads are attempted with
cached=FALSE
- Previously tried to access non-existent
Enhanced Dataset Registry
-
Added:
manual_updateflag todatasets.yamlfor FGV IBRE and NRE-IRE -
Added:
update_notesfield documenting why fresh downloads aren’t available -
Improved: Clear documentation in
_targets.Rexplaining data source choices
Files Changed
-
_targets.R: Updatedfetch_dataset()to supportsourceparameter; FGV and NRE-IRE now usesource="github" -
R/get_fgv_ibre.R: Removed broken internal data fallback; added clear error for fresh downloads -
R/get_nre_ire.R: Removed broken internal data fallback; added clear error for fresh downloads -
inst/extdata/datasets.yaml: Added manual update flags and notes
Bug Fixes from Recent Commits
Property Records Simplification (Commit 9eab0ca)
-
Refactored: Major simplification of
get_property_records.R(14% code reduction: 780→673 lines) -
Removed: Deprecated functions
get_ri_capitals()andget_ri_aggregates()with warning messages -
Removed: Unused metadata attributes (
source,download_time,download_info) that were never used - Simplified: Documentation for internal function (removed verbose examples and sections)
-
Improved:
scrape_registro_imoveis_links()with better connection cleanup and reduced complexity
BCB Dataset Critical Fixes (Commit bb580c8)
BCB Real Estate
- Fixed: CLI message serialization error in targets pipeline
-
Fixed: Compute
nrow()before CLI interpolation to avoid closure issues
BCB Series - Graceful Degradation (CRITICAL)
- Fixed: Replaced batch download with individual series downloads for better reliability
- Fixed: Now returns successful series even if some fail (e.g., 14/15 instead of 0/15)
-
Added: Per-series retry logic with exponential backoff using
purrr::possibly()pattern - Added: Clear warnings showing which series failed
-
Restored: Commented-out table filtering logic - now filters by
bcb_categorywhentablespecified -
Improved: Metadata-driven approach using
bcb_metadatadynamically (now downloads all 140 series, not just 15)
Get Dataset Infrastructure
-
Fixed: BCB Real Estate table filtering by category in
get-dataset.R -
Fixed: BCB Series table filtering by
bcb_category -
Added: Support for
table="all"invalidate_and_resolve_table()function - Fixed: Proper mapping of user-facing table names to internal Portuguese categories
Get Dataset Critical Fixes (Commit ce4768b)
CLI Message Scoping
-
Fixed: Added
.envir = parent.frame()tocli::cli_inform()calls incli_user()andcli_debug() - Fixed: “cannot coerce type ‘closure’ to vector of type ‘character’” error
- Affected: Previously failed for rppi_bis, property_records, and all functions using these helpers
FipeZap Data Quality
-
Fixed: Added
standardize_city_names()call after binding FipeZap data - Fixed: Now correctly shows “Brazil” instead of “Índice Fipezap” for national index
Property Records Table Extraction
-
Fixed: Added special handling for nested
property_recordsstructure inget-dataset.R - Fixed: Now returns single tibbles instead of nested lists
- Fixed: All tables (capitals, cities, aggregates, transfers) now work correctly
Testing Infrastructure
-
Added: Comprehensive integration test suite with 37 tests covering critical
get_dataset()functionality -
Added: Tests with
source="fresh"to catch real-world failures before production - Added: GitHub Actions CI workflow for weekly integration tests
-
Added: Manual testing script
tests/basic_checks.Rfor development
realestatebr 0.4.1
Bug Fixes
RPPI Individual Table Access
-
Fixed:
get_dataset("rppi", "ivgr")and other individual RPPI tables now work correctly - Fixed: Vignette build errors caused by RPPI table routing issues
-
Improved: Internal
get_rppi()function now supports all individual RPPI tables (fipezap, ivgr, igmi, iqa, iqaiw, ivar, secovi_sp) in addition to stacked tables (sale, rent, all)
realestatebr 0.4.0
Major Breaking Changes - API Consolidation
Unified Data Interface
This release implements a major breaking change that consolidates 15+ individual get_*() functions into a single, unified get_dataset() interface. This dramatically simplifies the package API while maintaining full functionality.
BREAKING CHANGE: All individual get_*() functions have been removed: - get_abecip_indicators(), get_abrainc_indicators(), get_rppi(), get_bcb_realestate(), etc. - Migration: Use get_dataset("dataset_name") instead
RPPI Code Simplification (Internal)
Major refactoring of RPPI functions for better maintainability: - 67% code reduction: 1579 lines → 519 lines (1060 lines removed) - Bug fix: FipeZap national index now correctly standardized to name_muni == "Brazil" - Shared helpers: Created rppi-helpers.R with common functions to eliminate duplication - Removed overhead: Eliminated unused stack parameter, cli_debug calls, and metadata attributes - Simplified documentation: Removed verbose sections (Progress Reporting, Error Handling, Examples) from internal functions - All functions now @keywords internal: Only get_dataset() is user-facing
Benefits: - Easier to maintain and debug - Faster execution (less overhead) - Consistent error handling across all indices - Bug fixes apply to all functions automatically
CBIC Dataset - Partial Release (Cement Only)
Note: In v0.4.0, the CBIC dataset is limited to cement tables only (validated data). Steel and PIM tables will be added in v0.4.1.
Available in v0.4.0: - cement_monthly_consumption - Monthly cement consumption by state - cement_annual_consumption - Annual cement consumption by region - cement_production_exports - Production, consumption, and export data - cement_monthly_production - Monthly cement production by state - cement_cub_prices - CUB cement prices by state
Coming in v0.4.1: - Steel prices and production data - PIM industrial production indices
# Works in v0.4.0
get_dataset("cbic") # Default: cement_monthly_consumption
get_dataset("cbic", "cement_cub_prices")
# Will error with informative message
get_dataset("cbic", "steel_prices") # Deferred to v0.4.1New Internal Architecture
-
Internal fetch functions: Created 12 new
fetch_*()functions with@keywords internal -
Registry-driven: All datasets managed through centralized
inst/extdata/datasets.yaml -
Hierarchical RPPI: Consolidated
rppiandrppi_indicesinto single hierarchical structure -
Consistent parameters: All internal functions use standardized
table,cached,quiet,max_retries
Simplified Public API
New unified interface:
# Get data from any dataset
data <- get_dataset("abecip") # Default table
data <- get_dataset("abecip", table = "sbpe") # Specific table
data <- get_dataset("rppi", table = "fipezap") # Hierarchical access
# Discover datasets
datasets <- list_datasets()
info <- get_dataset_info("rppi")Removed functions (now internal): - get_abecip_indicators() → get_dataset("abecip") - get_abrainc_indicators() → get_dataset("abrainc") - get_rppi() → get_dataset("rppi") - get_bcb_realestate() → get_dataset("bcb_realestate") - get_bcb_series() → get_dataset("bcb_series") - Plus 10 more functions
Enhanced Data Access
- Smart fallback: Auto fallback from GitHub cache → fresh download
-
Source control: Explicit
source = "cache"/"github"/"fresh"options - Better error messages: Detailed troubleshooting information
- Metadata preservation: All data includes source tracking and download info
Comprehensive Testing
-
New test suite:
test-internal-functions-0.4.0.Rwith 100 tests - Registry validation: Ensures all datasets have proper internal function mappings
- Parameter consistency: Validates all internal functions follow standard interface
- Hierarchical testing: Comprehensive RPPI access pattern validation
Migration Guide
For Existing Code (Breaking Changes)
# OLD (0.3.x) - Will no longer work
data <- get_abecip_indicators(table = "sbpe")
data <- get_rppi(table = "fipezap")
data <- get_bcb_realestate(table = "all")
# NEW (0.4.0) - Required migration
data <- get_dataset("abecip", table = "sbpe")
data <- get_dataset("rppi", table = "fipezap")
data <- get_dataset("bcb_realestate", table = "all")Dataset Name Mapping
| Old Function | New get_dataset() Name |
|---|---|
get_abecip_indicators() |
"abecip" |
get_abrainc_indicators() |
"abrainc" |
get_rppi() |
"rppi" |
get_bcb_realestate() |
"bcb_realestate" |
get_bcb_series() |
"bcb_series" |
get_rppi_bis() |
"rppi_bis" |
get_secovi() |
"secovi" |
get_fgv_indicators() |
"fgv_indicators" |
get_b3_stocks() |
"b3_stocks" |
get_nre_ire() |
"nre_ire" |
get_cbic_*() |
"cbic" |
get_itbi() |
"itbi" |
get_property_records() |
"registro" |
RPPI Consolidation
# OLD - Multiple functions
fipezap <- get_rppi_fipezap()
igmi <- get_rppi_igmi()
bis <- get_rppi_bis()
# NEW - Unified hierarchical access
fipezap <- get_dataset("rppi", table = "fipezap")
igmi <- get_dataset("rppi", table = "igmi")
bis <- get_dataset("rppi", table = "bis")Technical Implementation
Internal Architecture
-
12 internal fetch functions:
fetch_rppi(),fetch_abecip(), etc. -
Registry system: Complete mapping in
datasets.yaml -
Fallback mechanism:
get_from_internal_function()→get_from_legacy_function() -
NAMESPACE cleanup: Only exports
get_dataset(),list_datasets(), utilities
Backward Compatibility
- Phase 1: Internal functions call legacy functions for gradual transition
- Testing: Comprehensive test coverage ensures functionality preservation
- Error handling: Graceful degradation with informative error messages
This release represents a major architectural shift toward a unified, maintainable API. While it introduces breaking changes, the new interface is significantly simpler and more powerful.
Full Changelog: https://github.com/viniciusoike/realestatebr/compare/v0.3.0…v0.4.0
realestatebr 0.3.0
Major Features and Improvements
Phase 2: Data Pipeline Implementation Complete
- {targets} Pipeline Framework: Implemented comprehensive targets workflow for automated data processing and validation
- Automated Data Workflows: Added daily and weekly GitHub Actions workflows using the targets pipeline
- Data Validation Infrastructure: Added comprehensive validation rules and reporting for all datasets
- Pipeline Performance Monitoring: Added automated report generation and validation status tracking
Enhanced Data Processing
-
Targets Pipeline:
_targets.Rworkflow with automated dependency management and parallel processing - Validation System: Comprehensive data validation rules with automated quality checks
- Pipeline Helpers: Centralized helper functions for consistent data processing across all sources
- Report Generation: Automated pipeline status reports and data quality summaries
Improved Function Reliability
-
Error Handling: Enhanced error handling in
cache.Rwith better fallback mechanisms -
Function Fixes: Fixed parameter bugs in
get_abrainc_indicators()(category → table) -
Data Access: Improved
get_nre_ire()to use internal package data directly -
Internal Data: Updated
sysdata.rdawith latest processed datasets
Infrastructure Improvements
- Workflow Automation: Replaced single update workflow with focused daily/weekly pipelines
- Cache Management: Improved cache validation and fallback strategies
- Data Source Updates: Enhanced FGV data cleaning with improved formatting
-
Dependency Updates: Added
targetsandtarchetypesto package dependencies
Technical Implementation
Targets Pipeline Architecture
- Automated Processing: All datasets now processed through unified targets pipeline
- Quality Assurance: Built-in validation and quality checks for all data sources
- Performance Monitoring: Real-time pipeline status and error reporting
- Dependency Management: Automatic detection of data updates and re-processing
Enhanced Error Handling
- Graceful Degradation: Improved fallback mechanisms for failed data retrievals
- Better Diagnostics: Enhanced error messages and troubleshooting information
- Retry Logic: Smart retry mechanisms with exponential backoff
- Progress Reporting: Real-time progress updates during long-running operations
Migration Notes
For Existing Users
- All existing functions continue to work unchanged
- Enhanced reliability and performance with new pipeline backend
- Improved error messages and troubleshooting information
- Better cache management and fallback strategies
For Developers
- New targets pipeline provides foundation for custom data workflows
- Enhanced validation framework for quality assurance
- Standardized helper functions for consistent data processing
- Comprehensive pipeline documentation and examples
This release establishes the foundation for automated data processing and validation, setting the stage for Phase 3 implementation with large dataset support.
Full Changelog: https://github.com/viniciusoike/realestatebr/compare/v0.2.0…v0.3.0
realestatebr 0.2.0
Major Features and Improvements
Phase 1 Modernization Complete
-
Modernized 13 core
get_*functions with consistent APIs, CLI-based error handling, and progress reporting -
Standardized function signatures with
table,cached,quiet, andmax_retriesparameters - Robust error handling with retry logic, exponential backoff, and informative error messages
- Enhanced documentation with comprehensive examples and @section blocks
New Unified Data Architecture
-
list_datasets()- Discover available datasets with filtering capabilities -
get_dataset()- Unified data access function with intelligent fallback -
Registry system in
inst/extdata/datasets.yamlfor centralized dataset management - Improved caching with standardized cache location and validation
API Standardization
-
Introduced
tableparameter replacingcategoryacross all functions -
Backward compatibility maintained with deprecation warnings for
categoryparameter -
Consistent return types - single tibble by default, list when
table = "all" - Metadata attributes on all returned data with source tracking and download info
New Data Sources
-
CBIC construction materials data:
-
get_cbic_cement()- Cement consumption, production, and CUB prices -
get_cbic_steel()- Steel prices and production data -
get_cbic_pim()- Industrial production indices
-
- Enhanced RPPI suite with improved coordination and error handling
- Updated B3 stock data with standardized column names
Breaking Changes
Modernized Functions
Fully Modernized (13 functions)
-
get_abecip_indicators()- ABECIP real estate financing data -
get_abrainc_indicators()- ABRAINC launches and sales data -
get_b3_stocks()- B3 stock market data with improved column naming -
get_bcb_realestate()- Central Bank real estate credit data -
get_bcb_series()- BCB macroeconomic time series -
get_rppi_bis()- Bank for International Settlements RPPI data -
get_cbic_cement()- CBIC cement industry data (NEW) -
get_cbic_steel()- CBIC steel industry data (NEW) -
get_cbic_pim()- CBIC industrial production data (NEW) -
get_fgv_indicators()- FGV construction confidence indicators -
get_nre_ire()- Real Estate Index from NRE-Poli USP -
get_property_records()- Property registration data with robust Excel processing -
get_rppi()- Comprehensive RPPI coordinator with all sources -
get_secovi()- SECOVI-SP real estate data with parallel processing
Legacy Functions (Maintained)
-
get_rppi_bis()- Main function with modernized backend and single tibble returns -
get_itbi()andget_itbi_bhe()- Planned for Phase 3 (DuckDB integration)
Infrastructure Improvements
New Architecture Components
- Dataset registry system with YAML configuration
- Unified cache management with validation and fallback
- Translation framework for multilingual support
- Helper function ecosystem for robust web operations
Developer Experience
- Comprehensive test suite with 35+ tests covering all modernized functions
- Consistent documentation patterns with @section blocks and examples
-
CLI-based development workflow with
devtoolsintegration - GitHub Actions for automated testing and deployment
Migration Guide
For Existing Code
# Old (deprecated but still works)
data <- get_abecip_indicators(category = "all")
# New (recommended)
data <- get_abecip_indicators(table = "all")For New Code
# Discover available datasets
datasets <- list_datasets()
# Get data with unified interface
data <- get_dataset("abecip_indicators")
# Use modernized functions with progress
data <- get_abecip_indicators(table = "indicators", quiet = FALSE)Technical Details
Dependencies
-
Added:
clifor modern error handling and progress reporting -
Enhanced: Better integration with
dplyr,readr,httr, andrvest - Maintained: Full backward compatibility with existing dependencies
Performance
- Improved web scraping with intelligent retry logic
- Faster cache access with optimized file structure
- Better memory usage with streaming and lazy loading where appropriate
This release represents the completion of Phase 1 modernization, establishing a solid foundation for Phase 2 (data pipeline automation) and Phase 3 (large dataset support with DuckDB).
Full Changelog: https://github.com/viniciusoike/realestatebr/compare/v0.1.5…v0.2.0